{
  "id": 20747,
  "title": "Simple LB 0.23800 solution (Keras + VGG_16 pretrained)",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20747",
  "author_name": "Jiao Dong",
  "post_date": "2016-05-05T22:05:32.663000",
  "votes": 62,
  "comment_count": 164,
  "views": 42162,
  "content": "<p>I have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)</p>\n\n<p>I call it &quot;simple solution&quot; since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.</p>\n\n<p>Feel free to ask if you have any questions !</p>\n\n<p>The code is based on ZFTurbo's thread\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\">Keras Sample Code</a></p>\n\n<p>For a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. </p>\n\n<p>All my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)</p>\n\n<p>Some experience I gained from my experiments:</p>\n\n<ul>\n<li><p>Pre-trained model</p>\n\n<ul><li><p>In Keras, there are many pre-trained models available online, like <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">VGG-16</a> and <a href=\"https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\">VGG-19</a>. In Caffe there're also  <a href=\"https://github.com/KaimingHe/deep-residual-networks\">ResNet-50,101,152 by Kaiming He, MSRA</a>. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.</p></li>\n<li><p>Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.</p></li>\n<li>When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.</li>\n<li>From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization <a href=\"http://caffe.berkeleyvision.org/\">Caffe</a> would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)</li>\n<li>Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training &quot;quick and dirty&quot; models probably will not work very well.</li></ul></li>\n<li><p>During Training</p>\n\n<ul><li><p>Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease <strong>VERY QUICKLY</strong> within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.</p></li>\n<li><p>Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see <strong>a trend of decreasing validation loss with minor fluctuation</strong>. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\" alt=\"# of epochs is too small\" title>\n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\" alt=\"Much better # of epochs, keep an eye on the trend of validation loss\" title>\n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\" alt=\"Trend\" title></p></li></ul></li>\n<li><p>Some other small changes I made</p>\n\n<ul><li><p>Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.</p></li>\n<li><p>I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.</p></li>\n<li>Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.</li>\n<li>I tried using <a href=\"http://pjreddie.com/darknet/yolo/\">darknet</a> to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like &quot;if you see that person in back seat, then.....&quot; However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 118880,
      "postDate": "2016-05-05T22:05:32.663Z",
      "content": "<p>I have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)</p>\n\n<p>I call it &quot;simple solution&quot; since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.</p>\n\n<p>Feel free to ask if you have any questions !</p>\n\n<p>The code is based on ZFTurbo's thread\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\">Keras Sample Code</a></p>\n\n<p>For a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. </p>\n\n<p>All my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)</p>\n\n<p>Some experience I gained from my experiments:</p>\n\n<ul>\n<li><p>Pre-trained model</p>\n\n<ul><li><p>In Keras, there are many pre-trained models available online, like <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">VGG-16</a> and <a href=\"https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\">VGG-19</a>. In Caffe there're also  <a href=\"https://github.com/KaimingHe/deep-residual-networks\">ResNet-50,101,152 by Kaiming He, MSRA</a>. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.</p></li>\n<li><p>Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.</p></li>\n<li>When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.</li>\n<li>From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization <a href=\"http://caffe.berkeleyvision.org/\">Caffe</a> would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)</li>\n<li>Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training &quot;quick and dirty&quot; models probably will not work very well.</li></ul></li>\n<li><p>During Training</p>\n\n<ul><li><p>Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease <strong>VERY QUICKLY</strong> within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.</p></li>\n<li><p>Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see <strong>a trend of decreasing validation loss with minor fluctuation</strong>. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\" alt=\"# of epochs is too small\" title>\n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\" alt=\"Much better # of epochs, keep an eye on the trend of validation loss\" title>\n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\" alt=\"Trend\" title></p></li></ul></li>\n<li><p>Some other small changes I made</p>\n\n<ul><li><p>Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.</p></li>\n<li><p>I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.</p></li>\n<li>Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.</li>\n<li>I tried using <a href=\"http://pjreddie.com/darknet/yolo/\">darknet</a> to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like &quot;if you see that person in back seat, then.....&quot; However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.</li></ul></li>\n</ul>",
      "rawMarkdown": "I have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)\r\n\r\nI call it \"simple solution\" since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.\r\n\r\nFeel free to ask if you have any questions !\r\n\r\nThe code is based on ZFTurbo's thread\r\n[Keras Sample Code][1]\r\n\r\n\r\nFor a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. \r\n\r\nAll my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)\r\n\r\nSome experience I gained from my experiments:\r\n\r\n - Pre-trained model\r\n    - In Keras, there are many pre-trained models available online, like [VGG-16][2] and [VGG-19][3]. In Caffe there're also  [ResNet-50,101,152 by Kaiming He, MSRA][4]. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.\r\n \r\n    - Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\r\n    - When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.\r\n   - From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization [Caffe][5] would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)\r\n   - Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training \"quick and dirty\" models probably will not work very well.\r\n\r\n - During Training\r\n\r\n   - Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease **VERY QUICKLY** within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.\r\n\r\n   - Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see **a trend of decreasing validation loss with minor fluctuation**. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \r\n![# of epochs is too small][6]\r\n![Much better # of epochs, keep an eye on the trend of validation loss][7]\r\n![Trend][8]\r\n\r\n - Some other small changes I made\r\n\r\n   - Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.\r\n\r\n   - I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.\r\n   - Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.\r\n   - I tried using [darknet][9] to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like \"if you see that person in back seat, then.....\" However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\r\n  [2]: https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n  [3]: https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\r\n  [4]: https://github.com/KaimingHe/deep-residual-networks\r\n  [5]: http://caffe.berkeleyvision.org/\r\n  [6]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\r\n  [7]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\r\n  [8]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\r\n  [9]: http://pjreddie.com/darknet/yolo/",
      "votes": 62
    },
    {
      "id": 119383,
      "postDate": "2016-05-09T17:48:11.463Z",
      "content": "<p>[quote=Manuele Tamburrano;119377]</p>\n\n<p>thank you for clarifications, but probably you are remembering something wrong.\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?</p>\n\n<p>And still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons</p>\n\n<p>[/quote]</p>\n\n<p>That's the problem. It seems to be an issue with newer versions of Keras. See here <a href=\"https://github.com/fchollet/keras/issues/2371\">https://github.com/fchollet/keras/issues/2371</a>.\nI experienced the same thing with the script converging very slowly to a much higher than reported loss and having incorrect param numbers and then I changed the pop layer code to:</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>as suggested in the issue and now it seems to be converging a lot faster and with a lower loss.</p>",
      "rawMarkdown": "[quote=Manuele Tamburrano;119377]\r\n\r\nthank you for clarifications, but probably you are remembering something wrong.\r\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?\r\n\r\nAnd still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\r\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons\r\n\r\n[/quote]\r\n\r\nThat's the problem. It seems to be an issue with newer versions of Keras. See here https://github.com/fchollet/keras/issues/2371.\r\nI experienced the same thing with the script converging very slowly to a much higher than reported loss and having incorrect param numbers and then I changed the pop layer code to:\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\nas suggested in the issue and now it seems to be converging a lot faster and with a lower loss.",
      "votes": 11
    },
    {
      "id": 121891,
      "postDate": "2016-05-30T19:06:58.657Z",
      "content": "<p>I have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.</p>\n\n<p>I will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. </p>\n\n<p>I spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.</p>",
      "rawMarkdown": "I have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.\r\n\r\nI will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. \r\n\r\nI spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.\r\n\r\n",
      "votes": 10
    },
    {
      "id": 118886,
      "postDate": "2016-05-05T22:52:00.203Z",
      "content": "<p>In addition, I came across this online course while searching for information about CNN. I have been following it for a while, it's  the best introductory course of CNN online.</p>\n\n<p><a href=\"http://cs231n.stanford.edu/\">http://cs231n.stanford.edu/</a></p>",
      "rawMarkdown": "In addition, I came across this online course while searching for information about CNN. I have been following it for a while, it's  the best introductory course of CNN online.\r\n\r\n[http://cs231n.stanford.edu/][1]\r\n\r\n\r\n  [1]: http://cs231n.stanford.edu/",
      "votes": 8
    },
    {
      "id": 121127,
      "postDate": "2016-05-24T07:23:05.467Z",
      "content": "<p>I ran the script and got ~0.6LB with single model.</p>\n\n<p>Here is my guess why the author can get 0.2LB.</p>\n\n<p>According to the author, </p>\n\n<ul>\n<li>&quot;I have attached the original python code that got 0.23800 loss score, using 8 folds with 15 epochs.&quot;</li>\n<li>&quot;If you want to run 'fast' experiments, I got a LB 0.32640 with only 2 folds, 3 epoch each&quot;</li>\n</ul>\n\n<p>In main.py line 365, he splits drivers into train_drivers and test_drivers with KFold. However, when training the model in line 378, he uses the full training set. So it seems to me that the cross-validation part actually just trains the model on same data many times with different init. </p>\n\n<p>Maybe training more models with different init can improve LB score.</p>",
      "rawMarkdown": "I ran the script and got ~0.6LB with single model.\r\n\r\nHere is my guess why the author can get 0.2LB.\r\n\r\nAccording to the author, \r\n\r\n - \"I have attached the original python code that got 0.23800 loss score, using 8 folds with 15 epochs.\"\r\n - \"If you want to run 'fast' experiments, I got a LB 0.32640 with only 2 folds, 3 epoch each\"\r\n\r\nIn main.py line 365, he splits drivers into train_drivers and test_drivers with KFold. However, when training the model in line 378, he uses the full training set. So it seems to me that the cross-validation part actually just trains the model on same data many times with different init. \r\n\r\nMaybe training more models with different init can improve LB score.",
      "votes": 3
    },
    {
      "id": 119189,
      "postDate": "2016-05-07T21:39:26.627Z",
      "content": "<p>Good point, the file I used do have restricted use, so for people who is considering to participate this competition seriously you should be aware of licence issue. </p>\n\n<p>But there're still plenty of models released based on ImageNet that are unrestricted, like <a href=\"http://caffe.berkeleyvision.org/model_zoo.html#bvlc-model-license\">Caffe's Model Zoo BVLC Model license</a>  and <a href=\"https://github.com/KaimingHe/deep-residual-networks\">ResNet's Third-Party re-implementations (including Kaggle)</a> , if you insist using VGG-16 in keras, I just found admin's updated response <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread/119178#post119178\">Here at post #28</a></p>\n\n<p>I have already moved on to other projects since end of April, but at least I showed what loss score you can get with a particular pre-trained model. There are many other models you can try. From admin's reminder of copyright, please make sure you chose the right model without commercial restriction. GL, HF :)</p>\n\n<p>[quote=Wendy Kan;119155]</p>\n\n<p>Hi all, </p>\n\n<p>Someone in the community flagged the usage of <a href=\"https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md\">VGG-16</a> here. We looked into the license and found this in their disclaimer:</p>\n\n<blockquote>\n  <p>license: <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a>\n  (non-commercial use only)</p>\n</blockquote>\n\n<p>Since it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Good point, the file I used do have restricted use, so for people who is considering to participate this competition seriously you should be aware of licence issue. \r\n\r\nBut there're still plenty of models released based on ImageNet that are unrestricted, like [Caffe's Model Zoo BVLC Model license][1]  and [ResNet's Third-Party re-implementations (including Kaggle)][2] , if you insist using VGG-16 in keras, I just found admin's updated response [Here at post #28][3]\r\n\r\nI have already moved on to other projects since end of April, but at least I showed what loss score you can get with a particular pre-trained model. There are many other models you can try. From admin's reminder of copyright, please make sure you chose the right model without commercial restriction. GL, HF :)\r\n\r\n[quote=Wendy Kan;119155]\r\n\r\nHi all, \r\n\r\nSomeone in the community flagged the usage of [VGG-16][4] here. We looked into the license and found this in their disclaimer:\r\n\r\n> license: http://creativecommons.org/licenses/by-nc/4.0/\r\n> (non-commercial use only)\r\n\r\nSince it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. \r\n\r\n[/quote]\r\n\r\n\r\n  [1]: http://caffe.berkeleyvision.org/model_zoo.html#bvlc-model-license\r\n  [2]: https://github.com/KaimingHe/deep-residual-networks\r\n  [3]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread/119178#post119178\r\n  [4]: https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md",
      "votes": 3
    },
    {
      "id": 125242,
      "postDate": "2016-06-28T02:09:37.653Z",
      "content": "<p>there are a few ways to do.  It is a &quot;trial and error&quot; and see which would work best.\nGiven a pretrained CNN network, we wish to change the size of the filter and weights.</p>\n\n<p>E.g. for making them larger,  we can:</p>\n\n<ol>\n<li>pad new values with zeros   </li>\n<li><p>fill new values with random values. The magnitude of the random values are important.\n You need to ensure:</p>\n\n<ul><li><p>distribution (old_output) = distribution(new_output) </p></li>\n<li><p>where: old_output = function(old_weight, old_input) ,  new_output =function(new_weight, old_input), and  function = conv or inner product</p></li></ul></li>\n</ol>\n\n<p>Assuming Gaussian, you just need to ensure mean and std of old_output and new_output are the same. Hence the new  random values can be just random Gaussian noise of appropriate std.</p>\n\n<p>for 224x224 to 256x256, I think you can just pad with zero and try first. I works for me as well.</p>\n\n<p>For reference, refer to:</p>\n\n<ul>\n<li><a href=\"http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization\">http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization</a></li>\n<li><a href=\"http://arxiv.org/abs/1511.06422\">http://arxiv.org/abs/1511.06422</a></li>\n<li><a href=\"http://deepdish.io/2015/02/24/network-initialization/\">http://deepdish.io/2015/02/24/network-initialization/</a></li>\n</ul>\n\n<p>the key of training deep networks is to make sure signals can propagate forward and backward.\nThe weights values (and data values) cannot be too large or too small. If too large, it will grow infinitely large and leads to explosion (you will see #NAN in training loss). If too small, it will grow infinitely small, aka the problem of diminishing gradient.</p>\n\n<p>[quote=Polaris;125174]</p>\n\n<p>@Heng CherKeng\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\nWhat's the meaning of &quot;just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.&quot;\nand &quot;As for the last fully connect layers, yo can randomly initialised. &quot;</p>\n\n<p>Since  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "there are a few ways to do.  It is a \"trial and error\" and see which would work best.\r\nGiven a pretrained CNN network, we wish to change the size of the filter and weights.\r\n\r\nE.g. for making them larger,  we can:\r\n\r\n 1. pad new values with zeros   \r\n 2. fill new values with random values. The magnitude of the random values are important.\r\n     You need to ensure:\r\n\r\n - distribution (old_output) = distribution(new_output) \r\n   \r\n\r\n - where: old_output = function(old_weight, old_input) ,  new_output =function(new_weight, old_input), and  function = conv or inner product\r\n\r\nAssuming Gaussian, you just need to ensure mean and std of old_output and new_output are the same. Hence the new  random values can be just random Gaussian noise of appropriate std.\r\n\r\nfor 224x224 to 256x256, I think you can just pad with zero and try first. I works for me as well.\r\n\r\nFor reference, refer to:\r\n\r\n - http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization\r\n - http://arxiv.org/abs/1511.06422\r\n - http://deepdish.io/2015/02/24/network-initialization/\r\n\r\nthe key of training deep networks is to make sure signals can propagate forward and backward.\r\nThe weights values (and data values) cannot be too large or too small. If too large, it will grow infinitely large and leads to explosion (you will see #NAN in training loss). If too small, it will grow infinitely small, aka the problem of diminishing gradient.\r\n\r\n\r\n[quote=Polaris;125174]\r\n\r\n@Heng CherKeng\r\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\r\nWhat's the meaning of \"just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.\"\r\nand \"As for the last fully connect layers, yo can randomly initialised. \"\r\n\r\nSince  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning\r\n\r\n\r\n\r\n[/quote]\r\n",
      "votes": 4
    },
    {
      "id": 119077,
      "postDate": "2016-05-07T02:00:10.940Z",
      "content": "<p>When you train, could you try comment out the initializing with the pretrain vgg16.pkl and just use random initialization? by doing so, we know how much it really helps with the pretrained net.</p>",
      "rawMarkdown": "When you train, could you try comment out the initializing with the pretrain vgg16.pkl and just use random initialization? by doing so, we know how much it really helps with the pretrained net.",
      "votes": 3
    },
    {
      "id": 118999,
      "postDate": "2016-05-06T16:14:20.677Z",
      "content": "<p>@rcarson</p>\n\n<p>In my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \nVal_loss is not reliable due to split by image not driver.</p>\n\n<p>In addition, by printing model.summary(), number of weights of last layer is 10010.\nIt indicates layers are not connected properly. </p>",
      "rawMarkdown": "@rcarson\r\n\r\nIn my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \r\nVal_loss is not reliable due to split by image not driver.\r\n\r\nIn addition, by printing model.summary(), number of weights of last layer is 10010.\r\nIt indicates layers are not connected properly. \r\n",
      "votes": 3
    },
    {
      "id": 119102,
      "postDate": "2016-05-07T08:02:25.173Z",
      "content": "<p><strong>Abhijay Arora</strong>, If use this code &quot;as is&quot; it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.</p>\n\n<p><a href=\"http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\">http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory</a></p>",
      "rawMarkdown": "**Abhijay Arora**, If use this code \"as is\" it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.\r\n\r\nhttp://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\r\n\r\n",
      "votes": 4
    },
    {
      "id": 129608,
      "postDate": "2016-08-01T00:02:03.060Z",
      "content": "<p>[quote=Mike Kim;129605]</p>\n\n<p>The geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 \n[/quote]</p>\n\n<p>It should not be an issue, at least because with softmax output 0 prediction can not happen due to the fact that </p>\n\n<blockquote>\n  <p>exp[x] = 0 &lt;=&gt; x = - Infinity</p>\n</blockquote>\n\n<p>And -Infinity should not appear because </p>\n\n<blockquote>\n  <p>min(float number stored in computer) &gt; -Infinity</p>\n</blockquote>\n\n<p>Although I can imaging zero probability appearing due to some rounding.</p>\n\n<p>Just looked through some submissions for this competition, I see very small numbers, but I do not see any zeros.</p>",
      "rawMarkdown": "[quote=Mike Kim;129605]\r\n\r\nThe geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 \r\n[/quote]\r\n\r\nIt should not be an issue, at least because with softmax output 0 prediction can not happen due to the fact that \r\n\r\n> exp[x] = 0 <=> x = - Infinity\r\n\r\n\r\nAnd -Infinity should not appear because \r\n\r\n> min(float number stored in computer) > -Infinity\r\n\r\nAlthough I can imaging zero probability appearing due to some rounding.\r\n\r\nJust looked through some submissions for this competition, I see very small numbers, but I do not see any zeros.",
      "votes": 1
    },
    {
      "id": 129347,
      "postDate": "2016-07-29T00:18:40.353Z",
      "content": "<p>As I found out theano on GPU would only accelerate operations with float32  as stated here :\n<a href=\"http://deeplearning.net/software/theano/tutorial/using_gpu.html\">http://deeplearning.net/software/theano/tutorial/using_gpu.html</a></p>\n\n<blockquote>\n  <p>What Can Be Accelerated on the GPU\n  The performance characteristics will change as we continue to optimize our implementations, and vary from device to device, but to give a rough idea of what to expect right now:</p>\n  \n  <p>Only computations with float32 data-type can be accelerated. Better support for float64 is expected in upcoming hardware but float64 computations are still relatively slow (Jan 2010).</p>\n</blockquote>",
      "rawMarkdown": "As I found out theano on GPU would only accelerate operations with float32  as stated here :\r\nhttp://deeplearning.net/software/theano/tutorial/using_gpu.html\r\n\r\n> What Can Be Accelerated on the GPU\r\nThe performance characteristics will change as we continue to optimize our implementations, and vary from device to device, but to give a rough idea of what to expect right now:\r\n\r\n>Only computations with float32 data-type can be accelerated. Better support for float64 is expected in upcoming hardware but float64 computations are still relatively slow (Jan 2010).",
      "votes": 1
    },
    {
      "id": 129101,
      "postDate": "2016-07-26T17:45:08.410Z",
      "content": "<p>@tetemin:</p>\n\n<ol>\n<li>Adam definitely helps. </li>\n<li>TensorFlow is roughly twice slower than  Theano =&gt; if you change your backend at aws it will spin faster. </li>\n<li>You  are finetuning =&gt; learning rate should be really small. I use 1e-5, 1e-6.</li>\n</ol>",
      "rawMarkdown": "@tetemin:\r\n\r\n 1. Adam definitely helps. \r\n 2. TensorFlow is roughly twice slower than  Theano => if you change your backend at aws it will spin faster. \r\n 3. You  are finetuning => learning rate should be really small. I use 1e-5, 1e-6.",
      "votes": 1
    },
    {
      "id": 128799,
      "postDate": "2016-07-24T03:28:34.617Z",
      "content": "<p>Hm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?</p>",
      "rawMarkdown": "Hm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?",
      "votes": 1
    },
    {
      "id": 128790,
      "postDate": "2016-07-24T00:20:48.670Z",
      "content": "<p>@NelsonChen</p>\n\n<p>With SWAP partition, you can run 224x224 on a 16GB memory PC, with just a bit change on test data loading process. </p>\n\n<p>Each epoch is taking 900s under GTX 1070 &amp; CUDA 8.0, 4x faster than AWS g2 instance.</p>",
      "rawMarkdown": "@NelsonChen\r\n\r\nWith SWAP partition, you can run 224x224 on a 16GB memory PC, with just a bit change on test data loading process. \r\n\r\nEach epoch is taking 900s under GTX 1070 & CUDA 8.0, 4x faster than AWS g2 instance.",
      "votes": 1
    },
    {
      "id": 126416,
      "postDate": "2016-07-08T09:05:36.113Z",
      "content": "<p>@VZ\nIt is simple to resolve your problem. Please put 'train' and 'test' inside a new fold named 'imgs'. Done.</p>",
      "rawMarkdown": "@VZ\r\nIt is simple to resolve your problem. Please put 'train' and 'test' inside a new fold named 'imgs'. Done.",
      "votes": 1
    },
    {
      "id": 126154,
      "postDate": "2016-07-06T21:11:35.650Z",
      "content": "<p>I just tried for the first time your script, but I am getting this error on:\ntrain_data = train_data.transpose((0, 3, 1, 2))</p>\n\n<p>ValueError: axes don't match array</p>",
      "rawMarkdown": "I just tried for the first time your script, but I am getting this error on:\r\ntrain_data = train_data.transpose((0, 3, 1, 2))\r\n\r\nValueError: axes don't match array",
      "votes": 1
    },
    {
      "id": 125403,
      "postDate": "2016-06-29T07:02:07.567Z",
      "content": "<p>Thanks everyone for sharing knowledge and know-how.</p>\n\n<p>Hello @Jadiel, I had a quick look at your script and you should be able to improve the top_model by using the weights for the 2 Dense 4096 layers from the VGG16 weigths and also use some init like &quot;he_normal&quot; for the last Dense 10 layer</p>",
      "rawMarkdown": "Thanks everyone for sharing knowledge and know-how.\r\n\r\n\r\nHello @Jadiel, I had a quick look at your script and you should be able to improve the top_model by using the weights for the 2 Dense 4096 layers from the VGG16 weigths and also use some init like \"he_normal\" for the last Dense 10 layer",
      "votes": 1
    },
    {
      "id": 125330,
      "postDate": "2016-06-28T17:25:38.167Z",
      "content": "<p>You can use numpy.memmap and map a file to memory. It is pretty fast, and does not harm your training performance.</p>\n\n<p>[quote=JennyYu;125250]</p>\n\n<p>@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. </p>\n\n<p>I learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  </p>\n\n<p>Good luck to you.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "You can use numpy.memmap and map a file to memory. It is pretty fast, and does not harm your training performance.\r\n\r\n[quote=JennyYu;125250]\r\n\r\n@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. \r\n\r\nI learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  \r\n\r\nGood luck to you.\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 124668,
      "postDate": "2016-06-21T02:07:25.677Z",
      "content": "<p>From my own understanding of normalization, it applies when you have multiple features in numerical format, sometimes in different numerical range. You would like to know for a particular data point, how much feature X deviates from other data points for the same attribute X. Sometimes the numerical value of feature would mislead you, for example, without any normalization a feature value 200 might appear to be twice as significant as feature value 100 of another data point, but if your average value for that feature field is 10,000, they might in fact be nearly the same, equally trivial. So for color images people often divide each pixel channel by 255 (0xFF) to normalize each color channel, also make sure no color channel's value would dominate others. </p>\n\n<p>For vgg pre-trained models, it's a bit different. If you refer to the <a href=\"https://arxiv.org/pdf/1409.1556.pdf\"> Original VGG Net paper by Karen Simonyan &amp; Andrew Zisserman</a> , in section 2.1 ARCHITECTURE, &quot;During training, the input to our ConvNets is a fixed-size 224 &#215; 224 RGB image. The only preprocessing we do is subtracting the mean RGB value, computed on the training set, from each pixel&quot;, they trained vgg model from 14+ millions of images on <a href=\"http://www.image-net.org/\">ImageNet</a>, and those &quot;magic numbers&quot; are the mean pixel value for each channel based on images they used for training, serving as benchmarks for each channel.</p>\n\n<p>[quote=JennyYu;124581]</p>\n\n<p>Hi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:</p>\n\n<p>&quot;Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.&quot;</p>\n\n<p>I've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!</p>\n\n<p>Jenny</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "From my own understanding of normalization, it applies when you have multiple features in numerical format, sometimes in different numerical range. You would like to know for a particular data point, how much feature X deviates from other data points for the same attribute X. Sometimes the numerical value of feature would mislead you, for example, without any normalization a feature value 200 might appear to be twice as significant as feature value 100 of another data point, but if your average value for that feature field is 10,000, they might in fact be nearly the same, equally trivial. So for color images people often divide each pixel channel by 255 (0xFF) to normalize each color channel, also make sure no color channel's value would dominate others. \r\n\r\nFor vgg pre-trained models, it's a bit different. If you refer to the [ Original VGG Net paper by Karen Simonyan & Andrew Zisserman][1] , in section 2.1 ARCHITECTURE, \"During training, the input to our ConvNets is a fixed-size 224 × 224 RGB image. The only preprocessing we do is subtracting the mean RGB value, computed on the training set, from each pixel\", they trained vgg model from 14+ millions of images on [ImageNet][2], and those \"magic numbers\" are the mean pixel value for each channel based on images they used for training, serving as benchmarks for each channel.\r\n\r\n\r\n[quote=JennyYu;124581]\r\n\r\nHi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:\r\n\r\n\"Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\"\r\n\r\nI've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!\r\n\r\nJenny\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: https://arxiv.org/pdf/1409.1556.pdf\r\n  [2]: http://www.image-net.org/",
      "votes": 1
    },
    {
      "id": 124465,
      "postDate": "2016-06-18T22:14:38.650Z",
      "content": "<p>i am training a vgg with 16 layers. All layers are fine tuned.\nThe split of the train and test set are based on drivers.\nrandom 4 drivers are used as validation and remaining 22 are used for training.</p>\n\n<p>From various submissions i made, i find that:</p>\n\n<ol>\n<li><p>So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.</p></li>\n<li><p>For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.</p></li>\n</ol>\n\n<p>[quote=Ehsan;124440]</p>\n\n<p>Very interesting!\nDid you keep first 11 layers and train remaining layers?\nDid you split validation based on drivers or just random?</p>\n\n<p>Thank you</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "i am training a vgg with 16 layers. All layers are fine tuned.\r\nThe split of the train and test set are based on drivers.\r\nrandom 4 drivers are used as validation and remaining 22 are used for training.\r\n\r\nFrom various submissions i made, i find that:\r\n\r\n 1. So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.\r\n\r\n 2. For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.\r\n \r\n\r\n[quote=Ehsan;124440]\r\n\r\nVery interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 124118,
      "postDate": "2016-06-15T17:13:35.170Z",
      "content": "<p>The code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>",
      "rawMarkdown": "The code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))",
      "votes": 1
    },
    {
      "id": 123902,
      "postDate": "2016-06-14T12:12:48.167Z",
      "content": "<p>I see. Its nice that you have very good scores even without data augmentation. I got some improvements with my Torch implementation but I guess I'm still missing some important parts. </p>\n\n<p>Btw, the main reason for having higher log loss in the public leaderboard than validation loss is not the difference in the numbers of image but is the fact that you are not splitting train/val images based on drivers. I once did the same thing at the beginning and also got very low validation loss. Now I'm splitting train/val set based on drivers and the validation loss seems to be a good indicator of the public leaderboard score.</p>\n\n<p>[quote=Jiao Dong;123700]</p>\n\n<p>You are right, it's the public leaderboard score.</p>\n\n<p>Considering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.</p>\n\n<p>[quote=Duc Nguyen;123697]</p>\n\n<p>Are those the loss scores in your training, or scores in the public leaderboard? \nI guess the later since you had much lower validation loss. \nIn this case, I think I am missing something in my Torch code.\nThanks again.</p>\n\n<p>[quote=Jiao Dong;123690]</p>\n\n<p>At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.</p>\n\n<p>[quote=Duc Nguyen;123685]</p>\n\n<p>@Jiao Dong Thanks for sharing the code.\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I see. Its nice that you have very good scores even without data augmentation. I got some improvements with my Torch implementation but I guess I'm still missing some important parts. \r\n\r\n\r\nBtw, the main reason for having higher log loss in the public leaderboard than validation loss is not the difference in the numbers of image but is the fact that you are not splitting train/val images based on drivers. I once did the same thing at the beginning and also got very low validation loss. Now I'm splitting train/val set based on drivers and the validation loss seems to be a good indicator of the public leaderboard score.\r\n\r\n[quote=Jiao Dong;123700]\r\n\r\nYou are right, it's the public leaderboard score.\r\n\r\nConsidering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.\r\n\r\n[quote=Duc Nguyen;123697]\r\n\r\nAre those the loss scores in your training, or scores in the public leaderboard? \r\nI guess the later since you had much lower validation loss. \r\nIn this case, I think I am missing something in my Torch code.\r\nThanks again.\r\n\r\n[quote=Jiao Dong;123690]\r\n\r\nAt 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 123590,
      "postDate": "2016-06-12T22:59:57.133Z",
      "content": "<ul>\n<li>No data augmentation was used. The image was just simply resized to 224x224. Color image was used.</li>\n</ul>\n\n<p>Yes, exactly.</p>\n\n<ul>\n<li>The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\n  one fold served as validation set and the remaining 7 as training set. The validation set was used \n  to determine when to terminate the training.</li>\n</ul>\n\n<p>The final model is the average of 8 models. The training phase is a bit different from your description: for each fold, you went over your entire training dataset and generate a model file, then there's a random split of training dataset such that random 85% is used for training and 15% is for validating, between each fold, it's highly likely that they are using different subset of training data / validation data, with different ordering.</p>\n\n<p>[quote=Heng CherKeng;123554]</p>\n\n<p>Can I confirm if my understanding for the  LB 0.23800 solution is correct or not?</p>\n\n<ul>\n<li><p>No data augmentation was used. The image was just simply resized to 224x224.\n  Color image was used.</p></li>\n<li><p>The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\n  one fold served as validation set and the remaining 7 as training set. The validation set was used \n  to determine when to terminate the training.</p></li>\n</ul>\n\n<p>Thanks!</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "- No data augmentation was used. The image was just simply resized to 224x224. Color image was used.\r\n\r\nYes, exactly.\r\n\r\n - The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\r\n      one fold served as validation set and the remaining 7 as training set. The validation set was used \r\n      to determine when to terminate the training.\r\n\r\nThe final model is the average of 8 models. The training phase is a bit different from your description: for each fold, you went over your entire training dataset and generate a model file, then there's a random split of training dataset such that random 85% is used for training and 15% is for validating, between each fold, it's highly likely that they are using different subset of training data / validation data, with different ordering.\r\n\r\n[quote=Heng CherKeng;123554]\r\n\r\nCan I confirm if my understanding for the  LB 0.23800 solution is correct or not?\r\n\r\n -  No data augmentation was used. The image was just simply resized to 224x224.\r\n      Color image was used.\r\n \r\n - The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\r\n      one fold served as validation set and the remaining 7 as training set. The validation set was used \r\n      to determine when to terminate the training.\r\n\r\nThanks!\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 121907,
      "postDate": "2016-05-30T20:53:15.863Z",
      "content": "<p>I tried vgg-19 with very limited amount of time spent on it ( I remember it was like 2 days before deadline so I can't possibly finish a good run) and its learning convergence gave me nearly the same validation loss with 15 epochs. I would say it might not be much better, but not worse than vgg-16 as well.</p>\n\n<p>[quote=kuan chen;119657]</p>\n\n<p>@ Jiao Dong\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I tried vgg-19 with very limited amount of time spent on it ( I remember it was like 2 days before deadline so I can't possibly finish a good run) and its learning convergence gave me nearly the same validation loss with 15 epochs. I would say it might not be much better, but not worse than vgg-16 as well.\r\n\r\n[quote=kuan chen;119657]\r\n\r\n@ Jiao Dong\r\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 121862,
      "postDate": "2016-05-30T12:47:34.587Z",
      "content": "<p>@Jiao Dong\nThx for sharing your code on the forum! But I have a question about your code.\nYou said that you used colored 224x224 images, but in your code, the variable color_type = 1 in your model. And when I try to load the pre-trained VGG weights and retrain, the model sends out an error message stating that the shape of the model weights are not compatible, (64,1,3,3) with (64,3,3,3).\nIt would be really great if you can help me out here. Thx!</p>",
      "rawMarkdown": "@Jiao Dong\r\nThx for sharing your code on the forum! But I have a question about your code.\r\nYou said that you used colored 224x224 images, but in your code, the variable color_type = 1 in your model. And when I try to load the pre-trained VGG weights and retrain, the model sends out an error message stating that the shape of the model weights are not compatible, (64,1,3,3) with (64,3,3,3).\r\nIt would be really great if you can help me out here. Thx!\r\n\r\n",
      "votes": 1
    },
    {
      "id": 121137,
      "postDate": "2016-05-24T08:19:57.597Z",
      "content": "<p>I don't use this script. I voted down because it seemed that this script didn't produce ~0.2 LB score. This will cause an confusion. The correct information should be shared.</p>",
      "rawMarkdown": "I don't use this script. I voted down because it seemed that this script didn't produce ~0.2 LB score. This will cause an confusion. The correct information should be shared.",
      "votes": 1
    },
    {
      "id": 121065,
      "postDate": "2016-05-23T12:05:19.553Z",
      "content": "<p>[quote=anokas;119526]</p>\n\n<p>I still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?</p>\n\n<p>My epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here</p>\n\n<p>[/quote]</p>\n\n<p>With a Titan X + CUDNN 7.5 (Winograd convolution algorithm) + Theano 0.8.2 (see .theanorc Settings below) I can get to 550s per epoch.</p>\n\n<pre><code>[dnn.conv]                                       \nalgo_fwd = time_once\nalgo_bwd_data = time_once\nalgo_bwd_filter = time_once\n</code></pre>\n\n<p>Removing the ZeroPadding2D layer and using Convolution2D with border_mode='same' instead of 'valid' reduces the time per epoch to 430s.</p>",
      "rawMarkdown": "[quote=anokas;119526]\r\n\r\nI still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?\r\n\r\nMy epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here\r\n\r\n[/quote]\r\n\r\nWith a Titan X + CUDNN 7.5 (Winograd convolution algorithm) + Theano 0.8.2 (see .theanorc Settings below) I can get to 550s per epoch.\r\n\r\n    [dnn.conv]                                       \r\n    algo_fwd = time_once\r\n    algo_bwd_data = time_once\r\n    algo_bwd_filter = time_once\r\n\r\nRemoving the ZeroPadding2D layer and using Convolution2D with border_mode='same' instead of 'valid' reduces the time per epoch to 430s.",
      "votes": 1
    },
    {
      "id": 119526,
      "postDate": "2016-05-11T07:13:52.643Z",
      "content": "<p>I still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?</p>\n\n<p>My epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here</p>",
      "rawMarkdown": "I still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?\r\n\r\nMy epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here",
      "votes": 1
    },
    {
      "id": 119476,
      "postDate": "2016-05-10T15:32:35.980Z",
      "content": "<p>My run with 25 epochs and 3 KFolds with the fix above ended with these values of loss and accuracy:</p>\n\n<p>loss: 0.0017 - acc: 0.9994 - val_loss: 0.0168 - val_acc: 0.9949</p>\n\n<p>Final LB score is ~0.6</p>\n\n<p>Anyway the validation is splitted on images and not by drivers so the validation loss is inaccurate, I'll run the train again with data splitted by driver ids to check when overfit occurs</p>",
      "rawMarkdown": "My run with 25 epochs and 3 KFolds with the fix above ended with these values of loss and accuracy:\r\n\r\nloss: 0.0017 - acc: 0.9994 - val_loss: 0.0168 - val_acc: 0.9949\r\n\r\nFinal LB score is ~0.6\r\n\r\nAnyway the validation is splitted on images and not by drivers so the validation loss is inaccurate, I'll run the train again with data splitted by driver ids to check when overfit occurs",
      "votes": 1
    },
    {
      "id": 119370,
      "postDate": "2016-05-09T15:50:10.673Z",
      "content": "<p>yes, I tried to lower learning rate and some small adjustment but the loss keeps decreasing very slowly, I get ~1.3 after 15 epochs.</p>\n\n<p>What version of Keras are you running? The one from pip (it should be 1.0.2) or master compiled from source? If this is the case, what's your last commit hash?</p>\n\n<p>Seems like you are using an old version, because your logs show accuracy values, but you only set &quot;show_accuracy=True&quot; in the fit method, but in the last version you should use metrics=[&quot;accuracy&quot;] in compile method</p>\n\n<p>[quote=Jiao Dong;119366]</p>\n\n<p>There's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.</p>\n\n<p>From replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "yes, I tried to lower learning rate and some small adjustment but the loss keeps decreasing very slowly, I get ~1.3 after 15 epochs.\r\n\r\nWhat version of Keras are you running? The one from pip (it should be 1.0.2) or master compiled from source? If this is the case, what's your last commit hash?\r\n\r\nSeems like you are using an old version, because your logs show accuracy values, but you only set \"show_accuracy=True\" in the fit method, but in the last version you should use metrics=[\"accuracy\"] in compile method\r\n\r\n[quote=Jiao Dong;119366]\r\n\r\nThere's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.\r\n\r\nFrom replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?\r\n\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 119159,
      "postDate": "2016-05-07T17:48:38.157Z",
      "content": "<p>@Wendy, but in <a href=\"https://github.com/albertomontesg/keras-model-zoo/tree/master/models/VGG-16\">here</a>, it says <code>License: unrestricted use</code>. Does it mean that it is restricted in Caffe, but unrestricted in Keras?</p>\n\n<p>Forgive me if I asked dumb questions. If VGG-16 is not allowed, does it mean that we are not allowed to use the pretrained weights during training, but we can still use its architecture?  Or both are not allowed?</p>",
      "rawMarkdown": "@Wendy, but in [here][1], it says `License: unrestricted use`. Does it mean that it is restricted in Caffe, but unrestricted in Keras?\r\n\r\nForgive me if I asked dumb questions. If VGG-16 is not allowed, does it mean that we are not allowed to use the pretrained weights during training, but we can still use its architecture?  Or both are not allowed?\r\n\r\n  [1]: https://github.com/albertomontesg/keras-model-zoo/tree/master/models/VGG-16",
      "votes": 1
    },
    {
      "id": 119119,
      "postDate": "2016-05-07T12:30:28.413Z",
      "content": "<p>@Abhijay, if you change the image size, you can try the code I posted <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20466/vgg-16-keras/117026#post117026\">here</a>. Add the following code before fullly connected layer. </p>\n\n<pre><code>assert os.path.exists(weights_path), 'Model weights not found (see &quot;weights_path&quot; variable in script).'\nf = h5py.File(weights_path)\nfor k in range(f.attrs['nb_layers']):\nif k &gt;= len(model.layers):\n    # we don't look at the last (fully-connected) layers in the savefile\n    break\ng = f['layer_{}'.format(k)]\nweights = [g['param_{}'.format(p)] for p in range(g.attrs['nb_params'])]\nmodel.layers[k].set_weights(weights)\nf.close()\nprint('Model loaded.')\n</code></pre>\n\n<p>Hope it helps. </p>",
      "rawMarkdown": "@Abhijay, if you change the image size, you can try the code I posted [here][1]. Add the following code before fullly connected layer. \r\n\r\n    assert os.path.exists(weights_path), 'Model weights not found (see \"weights_path\" variable in script).'\r\n    f = h5py.File(weights_path)\r\n    for k in range(f.attrs['nb_layers']):\r\n    if k >= len(model.layers):\r\n        # we don't look at the last (fully-connected) layers in the savefile\r\n        break\r\n    g = f['layer_{}'.format(k)]\r\n    weights = [g['param_{}'.format(p)] for p in range(g.attrs['nb_params'])]\r\n    model.layers[k].set_weights(weights)\r\n    f.close()\r\n    print('Model loaded.')\r\n\r\nHope it helps. \r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20466/vgg-16-keras/117026#post117026",
      "votes": 1
    },
    {
      "id": 118954,
      "postDate": "2016-05-06T10:31:07.840Z",
      "content": "<p>is there a smaller pretrained models?</p>\n\n<p>looks like it's too big for my hardware too.</p>",
      "rawMarkdown": "is there a smaller pretrained models?\r\n\r\nlooks like it's too big for my hardware too.",
      "votes": 1
    },
    {
      "id": 118923,
      "postDate": "2016-05-06T05:48:21.343Z",
      "content": "<p>I don't remember =.=  </p>\n\n<p>My roommate had been using tensorflow recently as well with similar problem. From papers and documents he read, tensorflow originally works upon google's infrastructure therefore for the open-source version a lot of optimizations and features are still not implemented, but I would expect its performance issue to be resolved in later version. News I read last week also mentioned google's DeepMind officially announced to switch to tensorflow from Torch for their projects, I would consider the open-source version of tensorflow is still under active development and optimization. </p>\n\n<p>Anyways, this script has left plenty of things not yet implemented , like image localization (properly deal with the other person at back seat), changing CNN structure, locking weights, changing activation function, changing learning parameter for particular layer, etc. From what I read from authors of VGG and ResNet, it's also worth trying to combine multiple models to improve overall classification performance. Even with the same single vgg-16 model, there are a whole bunch of things you can do other than just taking average. </p>\n\n<p>Wish this script is helpful to be a good starting point for later experiments based on pre-trained models :)\n[quote=gauss256;118920]</p>\n\n<p>[quote=Jiao Dong;118880]...it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU[/quote]</p>\n\n<p>Is this with TensorFlow 0.8? There are supposed to be speed improvements in the latest version.</p>\n\n<p>Many thanks for posting your script, looking forward to trying it out!</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "\r\nI don't remember =.=  \r\n\r\nMy roommate had been using tensorflow recently as well with similar problem. From papers and documents he read, tensorflow originally works upon google's infrastructure therefore for the open-source version a lot of optimizations and features are still not implemented, but I would expect its performance issue to be resolved in later version. News I read last week also mentioned google's DeepMind officially announced to switch to tensorflow from Torch for their projects, I would consider the open-source version of tensorflow is still under active development and optimization. \r\n\r\nAnyways, this script has left plenty of things not yet implemented , like image localization (properly deal with the other person at back seat), changing CNN structure, locking weights, changing activation function, changing learning parameter for particular layer, etc. From what I read from authors of VGG and ResNet, it's also worth trying to combine multiple models to improve overall classification performance. Even with the same single vgg-16 model, there are a whole bunch of things you can do other than just taking average. \r\n\r\nWish this script is helpful to be a good starting point for later experiments based on pre-trained models :)\r\n[quote=gauss256;118920]\r\n\r\n[quote=Jiao Dong;118880]...it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU[/quote]\r\n\r\nIs this with TensorFlow 0.8? There are supposed to be speed improvements in the latest version.\r\n\r\nMany thanks for posting your script, looking forward to trying it out!\r\n\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 125289,
      "postDate": "2016-06-28T13:00:23.817Z",
      "content": "<p>the attachment shows the results for 224x224 and 256x256 and the ensemble of the two.</p>\n\n<p>[quote=Vladimir Iglovikov;125238]</p>\n\n<p>After some tweaks and modifications I got 0.19 at public LB.</p>\n\n<p>Did anyone try to use bigger image size and not 224x224 ?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "the attachment shows the results for 224x224 and 256x256 and the ensemble of the two.\r\n\r\n\r\n[quote=Vladimir Iglovikov;125238]\r\n\r\nAfter some tweaks and modifications I got 0.19 at public LB.\r\n\r\nDid anyone try to use bigger image size and not 224x224 ?\r\n\r\n[/quote]\r\n",
      "votes": 2
    },
    {
      "id": 125004,
      "postDate": "2016-06-24T13:44:47.490Z",
      "content": "<p>Below is my result using vgg16 and batch size 16. I don't think a batch size of 16 would stop convergence.</p>\n\n<p>[quote=Ehsan;124802]</p>\n\n<p>Thanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Below is my result using vgg16 and batch size 16. I don't think a batch size of 16 would stop convergence.\r\n\r\n[quote=Ehsan;124802]\r\n\r\nThanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: http://file:///C:/Users/jgao/Desktop/1413561361361513413.JPG",
      "votes": 2
    },
    {
      "id": 124422,
      "postDate": "2016-06-18T10:11:15.093Z",
      "content": "<p>@ Jiao Dong Thank for your code and method. I followed your idea and did the implementation in C/C++ myself. Here is my implementation:</p>\n\n<ol>\n<li>Use all train samples to finetune a VGG16</li>\n<li>Repeat for eight times:\n\n<ul><li>Finetune (1) using a subset of train samples. Use the remaining samples as validation to decide when   to stop training</li></ul></li>\n<li>Submission results is the average of eight results from (2)</li>\n</ol>\n\n<p>Final results on leader board is 0.27857.</p>\n\n<p>Here are some differences of my implementation compared to @ Jiao Dong's</p>\n\n<ul>\n<li>use data argumentation (scale shift and rotate)</li>\n<li>very short training iterations. Fine tunning in step (2) above is limited to 2 epoch.</li>\n</ul>\n\n<p>Here is break down of performances of each models and their combination and the confusion matrix of the model later. </p>\n\n<p>I note that leader board results is very dependent of the selection of the validation drivers and how to terminate the training. (I think this is due insufficient training data provided). With some tweaks, you can improve the leader board score to about 0.23 as reported by @ Jiao Dong.</p>\n\n<p>To go beyond 0.23, what i did is to increase input image size to 256x256. you can still use the vgg16 imageNet pretrained model for initialization. Finally I ensembled the results of 224x224 and 256x256. This gives about 0.20.</p>",
      "rawMarkdown": "@ Jiao Dong Thank for your code and method. I followed your idea and did the implementation in C/C++ myself. Here is my implementation:\r\n\r\n 1.  Use all train samples to finetune a VGG16\r\n 2.  Repeat for eight times:\r\n    - Finetune (1) using a subset of train samples. Use the remaining samples as validation to decide when   to stop training\r\n 3. Submission results is the average of eight results from (2)\r\n\r\nFinal results on leader board is 0.27857.\r\n\r\nHere are some differences of my implementation compared to @ Jiao Dong's\r\n\r\n -  use data argumentation (scale shift and rotate)\r\n -  very short training iterations. Fine tunning in step (2) above is limited to 2 epoch.\r\n\r\n\r\nHere is break down of performances of each models and their combination and the confusion matrix of the model later. \r\n\r\nI note that leader board results is very dependent of the selection of the validation drivers and how to terminate the training. (I think this is due insufficient training data provided). With some tweaks, you can improve the leader board score to about 0.23 as reported by @ Jiao Dong.\r\n\r\nTo go beyond 0.23, what i did is to increase input image size to 256x256. you can still use the vgg16 imageNet pretrained model for initialization. Finally I ensembled the results of 224x224 and 256x256. This gives about 0.20.\r\n\r\n\r\n\r\n",
      "votes": 2
    },
    {
      "id": 122366,
      "postDate": "2016-06-03T12:27:57.367Z",
      "content": "<p>for the newer version of Keras, you need a slightly different code to do model surgery, i.e. pop out the last layer and insert the new layer</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>[quote=Jiao Dong;121903]</p>\n\n<p>@Manuele Tamburrano</p>\n\n<p>I think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)</p>\n\n<p>Because as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.</p>\n\n<p>I did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "for the newer version of Keras, you need a slightly different code to do model surgery, i.e. pop out the last layer and insert the new layer\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\n\r\n\r\n[quote=Jiao Dong;121903]\r\n\r\n@Manuele Tamburrano\r\n\r\nI think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)\r\n\r\nBecause as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.\r\n\r\nI did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(\r\n\r\n\r\n\r\n[/quote]\r\n",
      "votes": 2
    },
    {
      "id": 121894,
      "postDate": "2016-05-30T19:15:23.057Z",
      "content": "<p>Thank you very much for sharing your code @Jiao Dong, you don't have to record the video. Sorry for posting this late, but we want to confirm with the exact same code we got 0.21 LB. It does have some variance from device to device and from run to run. But as long as it converges, the result is very stable.</p>",
      "rawMarkdown": "Thank you very much for sharing your code @Jiao Dong, you don't have to record the video. Sorry for posting this late, but we want to confirm with the exact same code we got 0.21 LB. It does have some variance from device to device and from run to run. But as long as it converges, the result is very stable.",
      "votes": 2
    },
    {
      "id": 119155,
      "postDate": "2016-05-07T17:40:25.703Z",
      "content": "<p>Hi all, </p>\n\n<p>Someone in the community flagged the usage of <a href=\"https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md\">VGG-16</a> here. We looked into the license and found this in their disclaimer:</p>\n\n<blockquote>\n  <p>license: <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a>\n  (non-commercial use only)</p>\n</blockquote>\n\n<p>Since it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. </p>",
      "rawMarkdown": "Hi all, \r\n\r\nSomeone in the community flagged the usage of [VGG-16][1] here. We looked into the license and found this in their disclaimer:\r\n\r\n> license: http://creativecommons.org/licenses/by-nc/4.0/\r\n> (non-commercial use only)\r\n\r\nSince it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. \r\n\r\n  [1]: https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md",
      "votes": 2
    },
    {
      "id": 119149,
      "postDate": "2016-05-07T16:40:17.600Z",
      "content": "<p>I posted about how to use a memory mapped file in this thread here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147</a></p>\n\n<p>I also tried to run the network on a CPU, because my GPU doesn't have enough memory. So far, I only have one thread that's working on the data. Estimated time to finish a single epoch for the first fold:</p>\n\n<p>95 days.</p>\n\n<p>LOL! :)</p>",
      "rawMarkdown": "I posted about how to use a memory mapped file in this thread here:\r\n\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147\r\n\r\n\r\nI also tried to run the network on a CPU, because my GPU doesn't have enough memory. So far, I only have one thread that's working on the data. Estimated time to finish a single epoch for the first fold:\r\n\r\n95 days.\r\n\r\nLOL! :)\r\n",
      "votes": 2
    },
    {
      "id": 118979,
      "postDate": "2016-05-06T14:22:23.693Z",
      "content": "<p>I think on these lines you actually lost the connection between drivers and images:</p>\n\n<pre><code>perm = permutation(len(train_target))\ntrain_data = train_data[perm]\ntrain_target = train_target[perm]\n</code></pre>",
      "rawMarkdown": "I think on these lines you actually lost the connection between drivers and images:\r\n\r\n    perm = permutation(len(train_target))\r\n    train_data = train_data[perm]\r\n    train_target = train_target[perm]",
      "votes": 2
    },
    {
      "id": 120764,
      "postDate": "2016-05-20T12:27:02.723Z",
      "content": "<p>I find the proposed solution and corresponding LB score suspicious. With a pretrained AlexNet or pretrained VGG you will not get below 0.4 LB score. I have a hard time seeing what was done differently from just a common VGG here to get near 0.2. Note that the leap between 0.4LB and 0.2LB is very high.</p>",
      "rawMarkdown": "I find the proposed solution and corresponding LB score suspicious. With a pretrained AlexNet or pretrained VGG you will not get below 0.4 LB score. I have a hard time seeing what was done differently from just a common VGG here to get near 0.2. Note that the leap between 0.4LB and 0.2LB is very high."
    },
    {
      "id": 118924,
      "postDate": "2016-05-06T05:51:29.750Z",
      "content": "<p>I'm impressed you pulled off this score without knowing about Karpathy's course.</p>",
      "rawMarkdown": "I'm impressed you pulled off this score without knowing about Karpathy's course.",
      "votes": 1
    },
    {
      "id": 119097,
      "postDate": "2016-05-07T07:04:57.090Z",
      "content": "<p>[quote=Jiao Dong;118880]</p>\n\n<p>I have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)</p>\n\n<p>I call it &quot;simple solution&quot; since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.</p>\n\n<p>Feel free to ask if you have any questions !</p>\n\n<p>The code is based on ZFTurbo's thread\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\">Keras Sample Code</a></p>\n\n<p>For a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. </p>\n\n<p>All my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)</p>\n\n<p>Some experience I gained from my experiments:</p>\n\n<ul>\n<li><p>Pre-trained model</p>\n\n<ul><li><p>In Keras, there are many pre-trained models available online, like <a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">VGG-16</a> and <a href=\"https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\">VGG-19</a>. In Caffe there're also  <a href=\"https://github.com/KaimingHe/deep-residual-networks\">ResNet-50,101,152 by Kaiming He, MSRA</a>. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.</p></li>\n<li><p>Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.</p></li>\n<li>When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.</li>\n<li>From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization <a href=\"http://caffe.berkeleyvision.org/\">Caffe</a> would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)</li>\n<li>Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training &quot;quick and dirty&quot; models probably will not work very well.</li></ul></li>\n<li><p>During Training</p>\n\n<ul><li><p>Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease <strong>VERY QUICKLY</strong> within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.</p></li>\n<li><p>Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see <strong>a trend of decreasing validation loss with minor fluctuation</strong>. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\" alt=\"# of epochs is too small\" title>\n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\" alt=\"Much better # of epochs, keep an eye on the trend of validation loss\" title>\n<img src=\"https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\" alt=\"Trend\" title></p></li></ul></li>\n<li><p>Some other small changes I made</p>\n\n<ul><li><p>Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.</p></li>\n<li><p>I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.</p></li>\n<li>Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.</li>\n<li>I tried using <a href=\"http://pjreddie.com/darknet/yolo/\">darknet</a> to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like &quot;if you see that person in back seat, then.....&quot; However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.</li></ul></li>\n</ul>\n\n<p>[/quote]</p>\n\n<p>Firstly, thank you for sharing your ideas!\nI tried running your script as is, but got a memory error. I think this is because the image size is too large(224 x224). I have been training on sizes 64 x 64. Any tips on how can I get 224 x 224 sized images to fit into my RAM (4 GB)?  What changes need to be made to existing code?\nThanks!</p>",
      "rawMarkdown": "[quote=Jiao Dong;118880]\r\n\r\nI have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)\r\n\r\nI call it \"simple solution\" since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.\r\n\r\nFeel free to ask if you have any questions !\r\n\r\nThe code is based on ZFTurbo's thread\r\n[Keras Sample Code][1]\r\n\r\n\r\nFor a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. \r\n\r\nAll my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)\r\n\r\nSome experience I gained from my experiments:\r\n\r\n - Pre-trained model\r\n    - In Keras, there are many pre-trained models available online, like [VGG-16][2] and [VGG-19][3]. In Caffe there're also  [ResNet-50,101,152 by Kaiming He, MSRA][4]. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.\r\n \r\n    - Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\r\n    - When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.\r\n   - From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization [Caffe][5] would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)\r\n   - Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training \"quick and dirty\" models probably will not work very well.\r\n\r\n - During Training\r\n\r\n   - Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease **VERY QUICKLY** within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.\r\n\r\n   - Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see **a trend of decreasing validation loss with minor fluctuation**. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \r\n![# of epochs is too small][6]\r\n![Much better # of epochs, keep an eye on the trend of validation loss][7]\r\n![Trend][8]\r\n\r\n - Some other small changes I made\r\n\r\n   - Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.\r\n\r\n   - I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.\r\n   - Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.\r\n   - I tried using [darknet][9] to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like \"if you see that person in back seat, then.....\" However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\r\n  [2]: https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n  [3]: https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\r\n  [4]: https://github.com/KaimingHe/deep-residual-networks\r\n  [5]: http://caffe.berkeleyvision.org/\r\n  [6]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\r\n  [7]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\r\n  [8]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\r\n  [9]: http://pjreddie.com/darknet/yolo/\r\n\r\n[/quote]\r\n\r\n\r\n\r\nFirstly, thank you for sharing your ideas!\r\nI tried running your script as is, but got a memory error. I think this is because the image size is too large(224 x224). I have been training on sizes 64 x 64. Any tips on how can I get 224 x 224 sized images to fit into my RAM (4 GB)?  What changes need to be made to existing code?\r\nThanks!",
      "votes": -3
    },
    {
      "id": 129638,
      "postDate": "2016-08-01T07:03:04.660Z",
      "content": "<p>Probably geometry mean should be done with following procedure:\n1) fix 0 -&gt; 0.00001 and 1 -&gt; 0.99999\n2) Apply geometry mean on new fixed data.</p>",
      "rawMarkdown": "Probably geometry mean should be done with following procedure:\r\n1) fix 0 -> 0.00001 and 1 -> 0.99999\r\n2) Apply geometry mean on new fixed data."
    },
    {
      "id": 129615,
      "postDate": "2016-08-01T01:41:36.620Z",
      "content": "<p>It seems you're correct regarding the Keras models. I checked the minimums and it looks something like: 2.393357e-35 rather than 0. I originally thought it was 0 because R's summary function rounds on the display by default.</p>",
      "rawMarkdown": "It seems you're correct regarding the Keras models. I checked the minimums and it looks something like: 2.393357e-35 rather than 0. I originally thought it was 0 because R's summary function rounds on the display by default."
    },
    {
      "id": 129605,
      "postDate": "2016-07-31T23:12:41.033Z",
      "content": "<p>The geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 automatically regardless of the other models (e.g. given 3 models, the geomean of 0.99,1,0 is still 0). Kaggle bounds logloss, but there's still a big penalty for guessing 0 and being wrong. </p>",
      "rawMarkdown": "The geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 automatically regardless of the other models (e.g. given 3 models, the geomean of 0.99,1,0 is still 0). Kaggle bounds logloss, but there's still a big penalty for guessing 0 and being wrong. "
    },
    {
      "id": 129560,
      "postDate": "2016-07-31T07:57:55.760Z",
      "content": "<p><strong>HeshamEraqi</strong>, in all my experiments geom was worse than mean.</p>",
      "rawMarkdown": "**HeshamEraqi**, in all my experiments geom was worse than mean."
    },
    {
      "id": 129555,
      "postDate": "2016-07-31T03:40:45.687Z",
      "content": "<p>[quote=HeshamEraqi;129551]</p>\n\n<p>Did anyone test<code>merge_several_folds_mean</code> versus <code>merge_several_folds_geom</code> on LB ?\nDo they give same LB as expected ?</p>\n\n<p>[/quote]\nI tried geom and get match worse score, you also have to renormalize after that as it no longer sums to 1 per row. </p>\n\n<p>I also found if you use softmax with temperature &gt;1 you can get better result on LB, but this seemt ot work only for non ensembled model. </p>",
      "rawMarkdown": "[quote=HeshamEraqi;129551]\r\n\r\nDid anyone test`merge_several_folds_mean` versus `merge_several_folds_geom` on LB ?\r\nDo they give same LB as expected ?\r\n\r\n\r\n[/quote]\r\nI tried geom and get match worse score, you also have to renormalize after that as it no longer sums to 1 per row. \r\n\r\nI also found if you use softmax with temperature >1 you can get better result on LB, but this seemt ot work only for non ensembled model. "
    },
    {
      "id": 129551,
      "postDate": "2016-07-31T01:58:36.900Z",
      "content": "<p>Did anyone test<code>merge_several_folds_mean</code> versus <code>merge_several_folds_geom</code> on LB ?\nDo they give same LB as expected ?</p>",
      "rawMarkdown": "Did anyone test`merge_several_folds_mean` versus `merge_several_folds_geom` on LB ?\r\nDo they give same LB as expected ?\r\n"
    },
    {
      "id": 129488,
      "postDate": "2016-07-30T10:28:07.860Z",
      "content": "<p>Getting 0.325 with mean of 13 folds. I attach code mainly taken from here and there.</p>",
      "rawMarkdown": "Getting 0.325 with mean of 13 folds. I attach code mainly taken from here and there."
    },
    {
      "id": 129470,
      "postDate": "2016-07-30T03:44:56.773Z",
      "content": "<p>Getting 0.36797 on LB with first fold and 0.23746 with mean of all 8 fold.  Do you have similar results? Usually ensemble helps just a bit...</p>",
      "rawMarkdown": "Getting 0.36797 on LB with first fold and 0.23746 with mean of all 8 fold.  Do you have similar results? Usually ensemble helps just a bit..."
    },
    {
      "id": 129448,
      "postDate": "2016-07-29T19:31:15.973Z",
      "content": "<p>[quote=Vladimir Iglovikov;129101]</p>\n\n<p>@tetemin:</p>\n\n<ol>\n<li>Adam definitely helps. </li>\n<li>TensorFlow is roughly twice slower than  Theano =&gt; if you change your backend at aws it will spin faster. </li>\n<li>You  are finetuning =&gt; learning rate should be really small. I use 1e-5, 1e-6.</li>\n</ol>\n\n<p>[/quote]</p>\n\n<p>Yes, thats interesting, that most papers suggest using SGD for fine-tuning, but Adam and Nadam(Adam version with netsterov momentum) works match better in my cases it gets to 95% in 3 iterations compared to 6 iterations of SGD.  Deep learning is the new field and some recommendations out there become obsolete once some one tires not to use them. </p>\n\n<p>I guess it only works with pretreated bottleneck classifier, haven't tried without.</p>",
      "rawMarkdown": "[quote=Vladimir Iglovikov;129101]\r\n\r\n@tetemin:\r\n\r\n 1. Adam definitely helps. \r\n 2. TensorFlow is roughly twice slower than  Theano => if you change your backend at aws it will spin faster. \r\n 3. You  are finetuning => learning rate should be really small. I use 1e-5, 1e-6.\r\n\r\n[/quote]\r\n\r\n\r\nYes, thats interesting, that most papers suggest using SGD for fine-tuning, but Adam and Nadam(Adam version with netsterov momentum) works match better in my cases it gets to 95% in 3 iterations compared to 6 iterations of SGD.  Deep learning is the new field and some recommendations out there become obsolete once some one tires not to use them. \r\n\r\nI guess it only works with pretreated bottleneck classifier, haven't tried without.\r\n"
    },
    {
      "id": 129326,
      "postDate": "2016-07-28T19:53:29.123Z",
      "content": "<p>Can somebody explain what advantage of using float32/float64 instead of uint8? For example here:</p>\n\n<pre><code>train_target = np_utils.to_categorical(train_target, 10)\ntrain_data = train_data.astype('float32')\n</code></pre>",
      "rawMarkdown": "Can somebody explain what advantage of using float32/float64 instead of uint8? For example here:\r\n\r\n    train_target = np_utils.to_categorical(train_target, 10)\r\n    train_data = train_data.astype('float32')"
    },
    {
      "id": 129193,
      "postDate": "2016-07-27T12:41:18.263Z",
      "content": "<p>[quote=Vladimir Iglovikov;129101]</p>\n\n<p>@tetemin:</p>\n\n<ol>\n<li>Adam definitely helps. </li>\n<li>TensorFlow is roughly twice slower than  Theano =&gt; if you change your backend at aws it will spin faster. </li>\n<li>You  are finetuning =&gt; learning rate should be really small. I use 1e-5, 1e-6.</li>\n</ol>\n\n<p>[/quote]</p>\n\n<p>Thanks Vladimir, trying with Adam now. I managed to get Tensorflow up to Theano speed by transposing the weights and using tf dimension ordering instead of th.</p>\n\n<p>Any other tips on how you managed to get this down to the 0.2-0.3 range of scores since I don't have that much time to experiment. I'm considering trying the following:</p>\n\n<ul>\n<li>Training the last fully connected layers with bottleneck results first for the classification problem first and then fine-tuning the whole network. Is that worth trying or did you get good results just from end-to-end fine-tuning from the beginning?</li>\n<li>Augmenting data by cropping out random patches, is this necessary to get a decent score?</li>\n</ul>",
      "rawMarkdown": "[quote=Vladimir Iglovikov;129101]\r\n\r\n@tetemin:\r\n\r\n 1. Adam definitely helps. \r\n 2. TensorFlow is roughly twice slower than  Theano => if you change your backend at aws it will spin faster. \r\n 3. You  are finetuning => learning rate should be really small. I use 1e-5, 1e-6.\r\n\r\n[/quote]\r\n\r\nThanks Vladimir, trying with Adam now. I managed to get Tensorflow up to Theano speed by transposing the weights and using tf dimension ordering instead of th.\r\n\r\nAny other tips on how you managed to get this down to the 0.2-0.3 range of scores since I don't have that much time to experiment. I'm considering trying the following:\r\n\r\n- Training the last fully connected layers with bottleneck results first for the classification problem first and then fine-tuning the whole network. Is that worth trying or did you get good results just from end-to-end fine-tuning from the beginning?\r\n- Augmenting data by cropping out random patches, is this necessary to get a decent score?\r\n"
    },
    {
      "id": 129136,
      "postDate": "2016-07-27T00:16:21.507Z",
      "content": "<p>[quote=NelsonChen;128799]</p>\n\n<p>Hm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?</p>\n\n<p>[/quote]</p>\n\n<p>I was stuck with exactly the same problems but have solved them by following the advice of @Ferris, ie changing the SWAP partition size.</p>\n\n<p>I found the following link <a href=\"http://askubuntu.com/questions/178712/how-to-increase-swap-space?noredirect=1&lq=1\">http://askubuntu.com/questions/178712/how-to-increase-swap-space?noredirect=1&amp;lq=1</a> helpful.</p>\n\n<p>I was getting out of memory errors with 16Gb of ram and an a 5GB swap file. Increasing the swap file to 32GB removes the memory problems. In order to resize my swap partition I had to boot from a recovery disk, resize the partition above the existing swap partition and then expand the swap.</p>\n\n<p>If, like me, you hadn't realised what or how to change size of SWAP partition then hopefully this will be of use!</p>\n\n<p>Sadly for me there's probably not sufficient time left now for me to run the models..! </p>",
      "rawMarkdown": "[quote=NelsonChen;128799]\r\n\r\nHm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?\r\n\r\n[/quote]\r\n\r\nI was stuck with exactly the same problems but have solved them by following the advice of @Ferris, ie changing the SWAP partition size.\r\n\r\nI found the following link http://askubuntu.com/questions/178712/how-to-increase-swap-space?noredirect=1&lq=1 helpful.\r\n\r\nI was getting out of memory errors with 16Gb of ram and an a 5GB swap file. Increasing the swap file to 32GB removes the memory problems. In order to resize my swap partition I had to boot from a recovery disk, resize the partition above the existing swap partition and then expand the swap.\r\n\r\nIf, like me, you hadn't realised what or how to change size of SWAP partition then hopefully this will be of use!\r\n\r\nSadly for me there's probably not sufficient time left now for me to run the models..! "
    },
    {
      "id": 129099,
      "postDate": "2016-07-26T17:17:21.353Z",
      "content": "<p>Thanks Jadiel, would you recomend using SGD or something like adam? Also are people using decay/momentum or just a static learning rate here? I've just reduced it to 0.0001 with SGD and it seems to actually be converging now in the first epoch.</p>\n\n<p>Also, has anyone trained on batches this small and if not am I missing something, my GPU has 4GB of RAM and can't fit anything larger than a batch of 8.</p>",
      "rawMarkdown": "Thanks Jadiel, would you recomend using SGD or something like adam? Also are people using decay/momentum or just a static learning rate here? I've just reduced it to 0.0001 with SGD and it seems to actually be converging now in the first epoch.\r\n\r\nAlso, has anyone trained on batches this small and if not am I missing something, my GPU has 4GB of RAM and can't fit anything larger than a batch of 8."
    },
    {
      "id": 129095,
      "postDate": "2016-07-26T17:11:01.657Z",
      "content": "<p>You need to reduce your learning rate.  With a smaller batch size the learning rate needs to also be smaller.</p>",
      "rawMarkdown": "You need to reduce your learning rate.  With a smaller batch size the learning rate needs to also be smaller."
    },
    {
      "id": 129093,
      "postDate": "2016-07-26T16:54:07.450Z",
      "content": "<p>Hi, i'm fine-tuning the vgg-16 model on an AWS g2 instance with keras &amp; tensorflow backend. I'm only able to use a batch size of 8 before running out of memory, each epoch is also taking around 8000s which seems a lot longer than anyone else here. After 1 epoch so far I get no convergence at all, my loss has been around 14.7 since the beginning.</p>\n\n<p>I have pre-processed the images correctly (converted from RGB to BGR and done the VGG mean subtraction) and also converted the Theano type model to a Tensorflow type.</p>\n\n<p>Does anyone have any advice, how have people managed to get this working and converging to as low as 0.3?</p>",
      "rawMarkdown": "Hi, i'm fine-tuning the vgg-16 model on an AWS g2 instance with keras & tensorflow backend. I'm only able to use a batch size of 8 before running out of memory, each epoch is also taking around 8000s which seems a lot longer than anyone else here. After 1 epoch so far I get no convergence at all, my loss has been around 14.7 since the beginning.\r\n\r\nI have pre-processed the images correctly (converted from RGB to BGR and done the VGG mean subtraction) and also converted the Theano type model to a Tensorflow type.\r\n\r\nDoes anyone have any advice, how have people managed to get this working and converging to as low as 0.3?"
    },
    {
      "id": 129028,
      "postDate": "2016-07-26T00:45:47.093Z",
      "content": "<p>Thanks for sharing the code and your insight!</p>",
      "rawMarkdown": "Thanks for sharing the code and your insight!"
    },
    {
      "id": 128803,
      "postDate": "2016-07-24T05:42:55.540Z",
      "content": "<p>It seems that on the g2 instance, theano is only using my GPU ram (which is limited to 4gbs) but not switching to my system ram when the memory runs out. Anyone know how to fix this? Thanks!</p>",
      "rawMarkdown": "It seems that on the g2 instance, theano is only using my GPU ram (which is limited to 4gbs) but not switching to my system ram when the memory runs out. Anyone know how to fix this? Thanks!"
    },
    {
      "id": 128787,
      "postDate": "2016-07-23T22:33:15.730Z",
      "content": "<p>Anyone that decided to use this script with smaller images sizes (64 x 64 or 128 x 128) manage to get decent scores (&lt; LB 0.8)? I have limited memory ~15 gb and I can only load the images at a smaller size to fit, but I'm not getting good scores (~LB 2). I messed around with ZFTurbo's original keras script and the best I could do was 0.8, so I was trying to go with the pre-trained model approach. Any help would be appreciated! Thanks!</p>",
      "rawMarkdown": "Anyone that decided to use this script with smaller images sizes (64 x 64 or 128 x 128) manage to get decent scores (< LB 0.8)? I have limited memory ~15 gb and I can only load the images at a smaller size to fit, but I'm not getting good scores (~LB 2). I messed around with ZFTurbo's original keras script and the best I could do was 0.8, so I was trying to go with the pre-trained model approach. Any help would be appreciated! Thanks!"
    },
    {
      "id": 128665,
      "postDate": "2016-07-22T12:46:29.247Z",
      "content": "<p>@Sandeep42\n I had similar problem while loading test data, and I solved it by loading and then predicting in several batches. Haven't figured out the workaround on training data yet. </p>",
      "rawMarkdown": "@Sandeep42\r\n I had similar problem while loading test data, and I solved it by loading and then predicting in several batches. Haven't figured out the workaround on training data yet. "
    },
    {
      "id": 128631,
      "postDate": "2016-07-22T04:25:37.097Z",
      "content": "<p>A silly question maybe, but I'm suspicious, does it make sense to do the following :</p>\n\n<pre><code>train_data = np.array(train_data, dtype=np.uint8)\n...\ntrain_data = train_data.astype('float32')\n</code></pre>\n\n<p>, instead of doing it at once like:</p>\n\n<pre><code>train_data = np.array(train_data, dtype=np.float32)\n</code></pre>",
      "rawMarkdown": "A silly question maybe, but I'm suspicious, does it make sense to do the following :\r\n\r\n    train_data = np.array(train_data, dtype=np.uint8)\r\n    ...\r\n    train_data = train_data.astype('float32')\r\n\r\n, instead of doing it at once like:\r\n\r\n    train_data = np.array(train_data, dtype=np.float32)"
    },
    {
      "id": 128616,
      "postDate": "2016-07-21T23:16:03.507Z",
      "content": "<p>[quote=Sandeep42;128547]</p>\n\n<p>[quote=HeshamEraqi;126450]</p>\n\n<p>@Ehsan &amp; @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.</p>\n\n<p>[/quote]</p>\n\n<p>I have been working on this problem from a couple of days. I have the same problem as you mentioned. It seems to me that when subtracting the mean pixel from the resized image, it creates a float. My RAM is a bottleneck when images are converted into floating points. Is there any way to get around this problem?</p>\n\n<p>[/quote]</p>\n\n<p>I believe that truncating the float's into uint8's won't highly deteriorate the performance.</p>\n\n<p>@Ferris: The script is OK, the float-int issue arose from some custom script I've added then I fixed it. </p>",
      "rawMarkdown": "[quote=Sandeep42;128547]\r\n\r\n[quote=HeshamEraqi;126450]\r\n\r\n@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.\r\n\r\n[/quote]\r\n\r\nI have been working on this problem from a couple of days. I have the same problem as you mentioned. It seems to me that when subtracting the mean pixel from the resized image, it creates a float. My RAM is a bottleneck when images are converted into floating points. Is there any way to get around this problem?\r\n\r\n[/quote]\r\n\r\nI believe that truncating the float's into uint8's won't highly deteriorate the performance.\r\n\r\n@Ferris: The script is OK, the float-int issue arose from some custom script I've added then I fixed it. \r\n\r\n"
    },
    {
      "id": 128600,
      "postDate": "2016-07-21T17:32:09.697Z",
      "content": "<p>I have two questions about pre-processing of the images for a VGG model. I appreciate it if you share your thoughts. For these questions, I use the python cv2 package to read images. These questions basically ask how the original VGG model was trained.</p>\n\n<ol>\n<li>I see two different approaches for changing the axis of color channels in the shared scripts. Approach (a) uses the transpose method: \nimg = img.transpose((2,0,1))\nand approach (b) uses the swapaxes:\nimg = img.swapaxes(2, 0)</li>\n</ol>\n\n<p>Approach (b) actually rotates the picture by 90 degrees as seen below:</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>img.shape</p>\n      \n      <p>(480, 640, 3)</p>\n      \n      <p>img.transpose((2, 0, 1)).shape</p>\n      \n      <p>(3, 480, 640)</p>\n      \n      <p>img.swapaxes(2, 0).shape</p>\n      \n      <p>(3, 640, 480)</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>Which approach is actually correct in the sense that VGG was trained using that? My feeling is that approach a in which the image is not rotated is correct but ironically I get worse results when I use that.</p>\n\n<ol start=\"2\">\n<li>As mentioned by @Ehsan, the cv2.imread method by default reads images in BGR format. So should we also change the color space (i.e.  img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)) if VGG is trained based on RGB.  For this part, the VGG paper explicitly states RGB but this color space transformation seems to be missing from @Jiao 's code. So I am not sure if it is necessary or not. Again, when I add the transformation from BRG to RGB to my code, the performance deteriorates. </li>\n</ol>\n\n<p>I also see the same deterioration in performance when I subtract the means from each color channel. Of course, the performance deterioration may be caused by overfitting but I am still wondering what is the proper approach  based on how the original model was trained.</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "I have two questions about pre-processing of the images for a VGG model. I appreciate it if you share your thoughts. For these questions, I use the python cv2 package to read images. These questions basically ask how the original VGG model was trained.\r\n\r\n1. I see two different approaches for changing the axis of color channels in the shared scripts. Approach (a) uses the transpose method: \r\n    img = img.transpose((2,0,1))\r\nand approach (b) uses the swapaxes:\r\n   img = img.swapaxes(2, 0)\r\n   \r\nApproach (b) actually rotates the picture by 90 degrees as seen below:\r\n\r\n>>> img.shape\r\n\r\n>>>(480, 640, 3)\r\n\r\n>>> img.transpose((2, 0, 1)).shape\r\n\r\n>>> (3, 480, 640)\r\n\r\n>>> img.swapaxes(2, 0).shape\r\n\r\n>>>(3, 640, 480)\r\n\r\nWhich approach is actually correct in the sense that VGG was trained using that? My feeling is that approach a in which the image is not rotated is correct but ironically I get worse results when I use that.\r\n\r\n2. As mentioned by @Ehsan, the cv2.imread method by default reads images in BGR format. So should we also change the color space (i.e.  img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)) if VGG is trained based on RGB.  For this part, the VGG paper explicitly states RGB but this color space transformation seems to be missing from @Jiao 's code. So I am not sure if it is necessary or not. Again, when I add the transformation from BRG to RGB to my code, the performance deteriorates. \r\n\r\nI also see the same deterioration in performance when I subtract the means from each color channel. Of course, the performance deterioration may be caused by overfitting but I am still wondering what is the proper approach  based on how the original model was trained.\r\n\r\nThanks.\r\n "
    },
    {
      "id": 128547,
      "postDate": "2016-07-21T04:52:32.037Z",
      "content": "<p>[quote=HeshamEraqi;126450]</p>\n\n<p>@Ehsan &amp; @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.</p>\n\n<p>[/quote]</p>\n\n<p>I have been working on this problem from a couple of days. I have the same problem as you mentioned. It seems to me that when subtracting the mean pixel from the resized image, it creates a float. My RAM is a bottleneck when images are converted into floating points. Is there any way to get around this problem?</p>",
      "rawMarkdown": "[quote=HeshamEraqi;126450]\r\n\r\n@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.\r\n\r\n[/quote]\r\n\r\nI have been working on this problem from a couple of days. I have the same problem as you mentioned. It seems to me that when subtracting the mean pixel from the resized image, it creates a float. My RAM is a bottleneck when images are converted into floating points. Is there any way to get around this problem?\r\n\r\n"
    },
    {
      "id": 128470,
      "postDate": "2016-07-20T11:02:22.573Z",
      "content": "<p>[quote=HeshamEraqi;126450]</p>\n\n<p>@Ehsan &amp; @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.</p>\n\n<p>[/quote]</p>\n\n<p>Hi HeshamEraqi,</p>\n\n<p>I think I'm having the same problem. Can you share a tip please? Thanks.</p>",
      "rawMarkdown": "[quote=HeshamEraqi;126450]\r\n\r\n@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.\r\n\r\n[/quote]\r\n\r\nHi HeshamEraqi,\r\n\r\nI think I'm having the same problem. Can you share a tip please? Thanks."
    },
    {
      "id": 128461,
      "postDate": "2016-07-20T10:10:32.727Z",
      "content": "<p>@RafayZiaMir\ndarknet/src/yolo.c function test_yolo</p>",
      "rawMarkdown": "@RafayZiaMir\r\ndarknet/src/yolo.c function test_yolo"
    },
    {
      "id": 128437,
      "postDate": "2016-07-20T06:37:51.260Z",
      "content": "<p>@ChrisJun very clear answer, thanks a lot.</p>",
      "rawMarkdown": "@ChrisJun very clear answer, thanks a lot."
    },
    {
      "id": 128222,
      "postDate": "2016-07-18T15:33:51.353Z",
      "content": "<p>Does anyone know which file(C file) we need to change in darknet to crop image? actually i cant find that file which i need to change</p>",
      "rawMarkdown": "Does anyone know which file(C file) we need to change in darknet to crop image? actually i cant find that file which i need to change"
    },
    {
      "id": 127357,
      "postDate": "2016-07-14T01:01:51.937Z",
      "content": "<p>@nanoix9</p>\n\n<p>mean_pixel comes from ImageNet pre-trained VGG16 model.</p>\n\n<p><a href=\"https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\">https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3</a></p>\n\n<p>In this link, you will see images are preprocessed in the order of </p>\n\n<p>1) Resize : Resizing image to 224x224 to fit the input shape of VGG16 model</p>\n\n<p>2) Subtract mean_pixel : Subtracting ImageNet mean pixels of RGB values to get zero-centered data(zero-centered data is almost always preferred in neural network)</p>\n\n<p>It is a default pre-processing values when using ImageNet pre-trained model.</p>\n\n<p>It comes from averaging the RGB(Red, Green, Blue) values of ImageNet data.</p>\n\n<p>If you want to train data from scratch, you should use mean_pixel of you training data.</p>\n\n<p>Chris</p>\n\n<p>[quote=nanoix9;127053]</p>\n\n<p>@Jiao Dong I saw this in the preprocessing part of code</p>\n\n<pre><code>mean_pixel = [103.939, 116.779, 123.68]\nfor c in range(3):\n    train_data[:, c, :, :] = train_data[:, c, :, :] - mean_pixel[c]\n</code></pre>\n\n<p>where is the <code>mean_pixel</code> comes from? Can I use some formula instead of hard coded numbers?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "@nanoix9\r\n\r\nmean_pixel comes from ImageNet pre-trained VGG16 model.\r\n\r\nhttps://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n\r\nIn this link, you will see images are preprocessed in the order of \r\n\r\n1) Resize : Resizing image to 224x224 to fit the input shape of VGG16 model\r\n\r\n2) Subtract mean_pixel : Subtracting ImageNet mean pixels of RGB values to get zero-centered data(zero-centered data is almost always preferred in neural network)\r\n\r\nIt is a default pre-processing values when using ImageNet pre-trained model.\r\n\r\nIt comes from averaging the RGB(Red, Green, Blue) values of ImageNet data.\r\n\r\n\r\nIf you want to train data from scratch, you should use mean_pixel of you training data.\r\n\r\nChris\r\n\r\n[quote=nanoix9;127053]\r\n\r\n@Jiao Dong I saw this in the preprocessing part of code\r\n\r\n    mean_pixel = [103.939, 116.779, 123.68]\r\n    for c in range(3):\r\n        train_data[:, c, :, :] = train_data[:, c, :, :] - mean_pixel[c]\r\n\r\nwhere is the `mean_pixel` comes from? Can I use some formula instead of hard coded numbers?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 127053,
      "postDate": "2016-07-13T06:29:03.410Z",
      "content": "<p>@Jiao Dong I saw this in the preprocessing part of code</p>\n\n<pre><code>mean_pixel = [103.939, 116.779, 123.68]\nfor c in range(3):\n    train_data[:, c, :, :] = train_data[:, c, :, :] - mean_pixel[c]\n</code></pre>\n\n<p>where is the <code>mean_pixel</code> comes from? Can I use some formula instead of hard coded numbers?</p>",
      "rawMarkdown": "@Jiao Dong I saw this in the preprocessing part of code\r\n\r\n    mean_pixel = [103.939, 116.779, 123.68]\r\n    for c in range(3):\r\n        train_data[:, c, :, :] = train_data[:, c, :, :] - mean_pixel[c]\r\n\r\nwhere is the `mean_pixel` comes from? Can I use some formula instead of hard coded numbers?"
    },
    {
      "id": 126452,
      "postDate": "2016-07-08T14:48:28.643Z",
      "content": "<p>Thanks for sharing. But i gave it up for my GPU is so weak. </p>",
      "rawMarkdown": "Thanks for sharing. But i gave it up for my GPU is so weak. "
    },
    {
      "id": 126450,
      "postDate": "2016-07-08T14:44:36.500Z",
      "content": "<p>@Ehsan &amp; @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.</p>",
      "rawMarkdown": "@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats."
    },
    {
      "id": 126374,
      "postDate": "2016-07-08T00:08:30.477Z",
      "content": "<p>@ HeshamEraqi\nAlso, try shuffle with different seed. It should converge in first epoch.</p>",
      "rawMarkdown": "@ HeshamEraqi\r\nAlso, try shuffle with different seed. It should converge in first epoch."
    },
    {
      "id": 126369,
      "postDate": "2016-07-07T23:22:56.317Z",
      "content": "<p>@ HeshamEraqi</p>\n\n<p>Try visualizing your image right before you feed into the keras model.\nAlso, try running your keras network in toy example (you can randomly download ~30 images from google).</p>\n\n<p>The point here is to identify whether the bug lies in the input image or model construction.\nYou learn the most through debugging after all :)</p>\n\n<p>Chris</p>\n\n<p>[quote=HeshamEraqi;126334]</p>\n\n<p>@xyz &amp; @Ehsan I tried decreasing learning rate, converted RGB to BGR, and scaled images to [0-1] nothing succeeds for me and loss is stuck ~2.3. Any hints what else could be the reason ?</p>\n\n<p>[quote=Ehsan;126113]</p>\n\n<p>Skimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].</p>\n\n<p>[quote=HeshamEraqi;126093]</p>\n\n<p>My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from <a href=\"https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\">here</a>. I use latest Keras version.\n<img src=\"https://s32.postimg.org/asm8jqqmt/eeee.png\" alt=\"enter image description here\" title></p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "@ HeshamEraqi\r\n\r\nTry visualizing your image right before you feed into the keras model.\r\nAlso, try running your keras network in toy example (you can randomly download ~30 images from google).\r\n\r\nThe point here is to identify whether the bug lies in the input image or model construction.\r\nYou learn the most through debugging after all :)\r\n\r\nChris\r\n\r\n[quote=HeshamEraqi;126334]\r\n\r\n@xyz & @Ehsan I tried decreasing learning rate, converted RGB to BGR, and scaled images to [0-1] nothing succeeds for me and loss is stuck ~2.3. Any hints what else could be the reason ?\r\n\r\n[quote=Ehsan;126113]\r\n\r\nSkimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 126334,
      "postDate": "2016-07-07T19:13:43.060Z",
      "content": "<p>@xyz &amp; @Ehsan I tried decreasing learning rate, converted RGB to BGR, and scaled images to [0-1] nothing succeeds for me and loss is stuck ~2.3. Any hints what else could be the reason ?</p>\n\n<p>[quote=Ehsan;126113]</p>\n\n<p>Skimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].</p>\n\n<p>[quote=HeshamEraqi;126093]</p>\n\n<p>My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from <a href=\"https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\">here</a>. I use latest Keras version.\n<img src=\"https://s32.postimg.org/asm8jqqmt/eeee.png\" alt=\"enter image description here\" title></p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "@xyz & @Ehsan I tried decreasing learning rate, converted RGB to BGR, and scaled images to [0-1] nothing succeeds for me and loss is stuck ~2.3. Any hints what else could be the reason ?\r\n\r\n[quote=Ehsan;126113]\r\n\r\nSkimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 126115,
      "postDate": "2016-07-06T11:33:40.007Z",
      "content": "<p>@Ehsan\nHi, Ehsan. Thank you for your reply. Could you mail me your email address? I have some questions and need for your help.\nThanks!</p>",
      "rawMarkdown": "@Ehsan\r\nHi, Ehsan. Thank you for your reply. Could you mail me your email address? I have some questions and need for your help.\r\nThanks!"
    },
    {
      "id": 126113,
      "postDate": "2016-07-06T11:17:23.960Z",
      "content": "<p>Skimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].</p>\n\n<p>[quote=HeshamEraqi;126093]</p>\n\n<p>My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from <a href=\"https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\">here</a>. I use latest Keras version.\n<img src=\"https://s32.postimg.org/asm8jqqmt/eeee.png\" alt=\"enter image description here\" title></p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Skimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 126097,
      "postDate": "2016-07-06T05:30:49.987Z",
      "content": "<p>@HeshamEraqi, you may need to decrease the learning rate and try again. </p>\n\n<p>[quote=HeshamEraqi;126093]</p>\n\n<p>My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from <a href=\"https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\">here</a>. I use latest Keras version.\n<img src=\"https://s32.postimg.org/asm8jqqmt/eeee.png\" alt=\"enter image description here\" title></p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "@HeshamEraqi, you may need to decrease the learning rate and try again. \r\n\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 126093,
      "postDate": "2016-07-06T04:39:35.130Z",
      "content": "<p>My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from <a href=\"https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\">here</a>. I use latest Keras version.\n<img src=\"https://s32.postimg.org/asm8jqqmt/eeee.png\" alt=\"enter image description here\" title></p>",
      "rawMarkdown": "My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png"
    },
    {
      "id": 125843,
      "postDate": "2016-07-03T11:24:01.367Z",
      "content": "<p>Here is my update on best single model\nSingle googlenet-CAM can give LB 0.38746 and single VGG-CAM give LB 0.27369. </p>\n\n<p>see below:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125842#post125842\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125842#post125842</a></p>",
      "rawMarkdown": "Here is my update on best single model\r\nSingle googlenet-CAM can give LB 0.38746 and single VGG-CAM give LB 0.27369. \r\n\r\nsee below:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125842#post125842"
    },
    {
      "id": 125690,
      "postDate": "2016-07-01T15:52:34.593Z",
      "content": "<p>I am sorry, it is a typo error. It should be 0.32. (and not 0.23)</p>\n\n<p>[quote=DavidGbodiOdaibo;125688]</p>\n\n<p>@Heng CherKeng, your vgg16 score is suspicious for a single network. If you are using the solution on this thread, your vgg16 score is based on an ensemble of 8 VGG16 models. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I am sorry, it is a typo error. It should be 0.32. (and not 0.23)\r\n\r\n[quote=DavidGbodiOdaibo;125688]\r\n\r\n@Heng CherKeng, your vgg16 score is suspicious for a single network. If you are using the solution on this thread, your vgg16 score is based on an ensemble of 8 VGG16 models. \r\n\r\n[/quote]\r\n"
    },
    {
      "id": 125688,
      "postDate": "2016-07-01T15:45:29.907Z",
      "content": "<p>@Heng CherKeng, your vgg16 score is suspicious for a single network. If you are using the solution on this thread, your vgg16 score is based on an ensemble of 8 VGG16 models. </p>",
      "rawMarkdown": "@Heng CherKeng, your vgg16 score is suspicious for a single network. If you are using the solution on this thread, your vgg16 score is based on an ensemble of 8 VGG16 models. "
    },
    {
      "id": 125680,
      "postDate": "2016-07-01T15:05:47.510Z",
      "content": "<p>Here is my best results for single network. I use finetunning from imageNet pretrained network.</p>\n\n<ul>\n<li>googlenet: 0.55</li>\n<li>vgg16: 0.32 </li>\n<li>resnet-50: cannot get it to work</li>\n</ul>\n\n<p>[quote=Ehsan;125674]</p>\n\n<p>Does anybody has a good result with VGG-19 or google-net or other networks?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Here is my best results for single network. I use finetunning from imageNet pretrained network.\r\n\r\n - googlenet: 0.55\r\n - vgg16: 0.32 \r\n - resnet-50: cannot get it to work\r\n\r\n[quote=Ehsan;125674]\r\n\r\nDoes anybody has a good result with VGG-19 or google-net or other networks?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 125679,
      "postDate": "2016-07-01T15:05:23.423Z",
      "content": "<p>I have tried fine-tuning VGGnet &amp; resnet in Torch. \nGot more or less the same results as you got. \nI expected resnet to bring better results but it was not the case.</p>\n\n<p>[quote=SecondPlan;125678]</p>\n\n<p>vgg-19: 0.20</p>\n\n<p>resnet:0.43</p>\n\n<p>googlenet:0.53</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I have tried fine-tuning VGGnet & resnet in Torch. \r\nGot more or less the same results as you got. \r\nI expected resnet to bring better results but it was not the case.\r\n\r\n[quote=SecondPlan;125678]\r\n\r\nvgg-19: 0.20\r\n\r\nresnet:0.43\r\n\r\ngooglenet:0.53\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 125678,
      "postDate": "2016-07-01T15:01:07.247Z",
      "content": "<p>vgg-19: 0.20</p>\n\n<p>resnet:0.43</p>\n\n<p>googlenet:0.53</p>",
      "rawMarkdown": "vgg-19: 0.20\r\n\r\nresnet:0.43\r\n\r\ngooglenet:0.53"
    },
    {
      "id": 125674,
      "postDate": "2016-07-01T14:46:49.177Z",
      "content": "<p>Does anybody has a good result with VGG-19 or google-net or other networks?</p>",
      "rawMarkdown": "Does anybody has a good result with VGG-19 or google-net or other networks?"
    },
    {
      "id": 125488,
      "postDate": "2016-06-29T22:22:15.430Z",
      "content": "<p>[quote=BenediktSchifferer;125427]</p>\n\n<p>I have the same problem like Jadiel.\nMy accuracy keeps at 0.10. But I use tensorflow with the pretrained model Inception_v3 from google.\nI may dont understand the way using pretrained models:</p>\n\n<ul>\n<li>I load the NN structure + load the weights of the pretrained model</li>\n<li>I delete the last layer (from last features -&gt; classes)</li>\n<li>I add my own classifier (from last features -&gt; my new classes)</li>\n<li>I process the image with the pretrained model until the last layer (so I get the feature vector)</li>\n<li>I train my classifier from features -&gt; my new classes</li>\n</ul>\n\n<p>Is this correct? Or do you load the weights and retrain the whole model over all layers?</p>\n\n<p>[/quote]</p>\n\n<p>after the third step, you need to train the model and then do your 4th step ( the features you obtain is finetuned towards your dataset)</p>",
      "rawMarkdown": "[quote=BenediktSchifferer;125427]\r\n\r\nI have the same problem like Jadiel.\r\nMy accuracy keeps at 0.10. But I use tensorflow with the pretrained model Inception_v3 from google.\r\nI may dont understand the way using pretrained models:\r\n\r\n* I load the NN structure + load the weights of the pretrained model\r\n* I delete the last layer (from last features -> classes)\r\n* I add my own classifier (from last features -> my new classes)\r\n* I process the image with the pretrained model until the last layer (so I get the feature vector)\r\n* I train my classifier from features -> my new classes\r\n\r\nIs this correct? Or do you load the weights and retrain the whole model over all layers?\r\n\r\n\r\n[/quote]\r\n\r\nafter the third step, you need to train the model and then do your 4th step ( the features you obtain is finetuned towards your dataset)\r\n"
    },
    {
      "id": 125427,
      "postDate": "2016-06-29T11:28:42.707Z",
      "content": "<p>I have the same problem like Jadiel.\nMy accuracy keeps at 0.10. But I use tensorflow with the pretrained model Inception_v3 from google.\nI may dont understand the way using pretrained models:</p>\n\n<ul>\n<li>I load the NN structure + load the weights of the pretrained model</li>\n<li>I delete the last layer (from last features -&gt; classes)</li>\n<li>I add my own classifier (from last features -&gt; my new classes)</li>\n<li>I process the image with the pretrained model until the last layer (so I get the feature vector)</li>\n<li>I train my classifier from features -&gt; my new classes</li>\n</ul>\n\n<p>Is this correct? Or do you load the weights and retrain the whole model over all layers?</p>",
      "rawMarkdown": "I have the same problem like Jadiel.\r\nMy accuracy keeps at 0.10. But I use tensorflow with the pretrained model Inception_v3 from google.\r\nI may dont understand the way using pretrained models:\r\n\r\n* I load the NN structure + load the weights of the pretrained model\r\n* I delete the last layer (from last features -> classes)\r\n* I add my own classifier (from last features -> my new classes)\r\n* I process the image with the pretrained model until the last layer (so I get the feature vector)\r\n* I train my classifier from features -> my new classes\r\n\r\nIs this correct? Or do you load the weights and retrain the whole model over all layers?\r\n"
    },
    {
      "id": 125362,
      "postDate": "2016-06-28T22:45:42.090Z",
      "content": "<p>It computes bottlenecks and uses generators to be gentle with memory. Unfortunately, my script does not work well.  The loss do not diminishes with time.  It's accuracy keeps at 0.10 after 7 epochs.  What could be wrong with it? Is it because I don't use batches? Is it that I am training wrong? Anyone has an idea?</p>",
      "rawMarkdown": "It computes bottlenecks and uses generators to be gentle with memory. Unfortunately, my script does not work well.  The loss do not diminishes with time.  It's accuracy keeps at 0.10 after 7 epochs.  What could be wrong with it? Is it because I don't use batches? Is it that I am training wrong? Anyone has an idea?"
    },
    {
      "id": 125250,
      "postDate": "2016-06-28T04:11:05.353Z",
      "content": "<p>@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. </p>\n\n<p>I learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  </p>\n\n<p>Good luck to you.</p>",
      "rawMarkdown": "@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. \r\n\r\nI learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  \r\n\r\nGood luck to you."
    },
    {
      "id": 125243,
      "postDate": "2016-06-28T02:27:08.750Z",
      "content": "<p>@Heng CherKeng\nOnce more,I sincerely thank you for your reply,that's meaningful and instructive.\nI will read over your words and references.\nIt is a &quot;trial and error&quot;,thank you!</p>",
      "rawMarkdown": "@Heng CherKeng\r\nOnce more,I sincerely thank you for your reply,that's meaningful and instructive.\r\nI will read over your words and references.\r\nIt is a \"trial and error\",thank you!"
    },
    {
      "id": 125241,
      "postDate": "2016-06-28T02:05:40.477Z",
      "content": "<p>@Vladimir Iglovikov\nI tried,but failed.\nCould you share the modifications you have done,such as image size and so on. \nThanks.\n[quote=Vladimir Iglovikov;125238]</p>\n\n<p>After some tweaks and modifications I got 0.19 at public LB.</p>\n\n<p>Did anyone try to use bigger image size and not 224x224 ?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "@Vladimir Iglovikov\r\nI tried,but failed.\r\nCould you share the modifications you have done,such as image size and so on. \r\nThanks.\r\n[quote=Vladimir Iglovikov;125238]\r\n\r\nAfter some tweaks and modifications I got 0.19 at public LB.\r\n\r\nDid anyone try to use bigger image size and not 224x224 ?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 125240,
      "postDate": "2016-06-28T01:54:45.253Z",
      "content": "<p>@JennyYu\nYeah, I tried LR e-2, -3,even -6, however it seems useless that the training loss still stunned at 14 after one epoch.\nI know something I understand in a wrong way,but I just have no idea about it.\nDid you solved your issues after you tried different learning rate?</p>",
      "rawMarkdown": "@JennyYu\r\nYeah, I tried LR e-2, -3,even -6, however it seems useless that the training loss still stunned at 14 after one epoch.\r\nI know something I understand in a wrong way,but I just have no idea about it.\r\nDid you solved your issues after you tried different learning rate?"
    },
    {
      "id": 125238,
      "postDate": "2016-06-28T01:32:35.710Z",
      "content": "<p>After some tweaks and modifications I got 0.19 at public LB.</p>\n\n<p>Did anyone try to use bigger image size and not 224x224 ?</p>",
      "rawMarkdown": "After some tweaks and modifications I got 0.19 at public LB.\r\n\r\nDid anyone try to use bigger image size and not 224x224 ?"
    },
    {
      "id": 125207,
      "postDate": "2016-06-27T16:48:42.250Z",
      "content": "<p>Thanks Jiao Dong for the script. I've been able to get a 0.32 LB score with eight folds and 5 epochs. </p>\n\n<p>I tried reproducing the VGG16-Keras results with Caffe (using the same configuration), but was unsuccessful. I tried to pretrain   GoogleNet and AlexNet as well  ... but did not get any decent results on Caffe.  Has anyone had any success with this problem using Caffe? </p>",
      "rawMarkdown": "Thanks Jiao Dong for the script. I've been able to get a 0.32 LB score with eight folds and 5 epochs. \r\n\r\nI tried reproducing the VGG16-Keras results with Caffe (using the same configuration), but was unsuccessful. I tried to pretrain   GoogleNet and AlexNet as well  ... but did not get any decent results on Caffe.  Has anyone had any success with this problem using Caffe? \r\n"
    },
    {
      "id": 125202,
      "postDate": "2016-06-27T15:28:02.120Z",
      "content": "<p>@ Polaris, I had similar issue as you. I used image size of 64x64 due to memory issues, and I randomly initialized the fully connected layers. My training loss was stunned due to learning rate, not because of my model. Did you try to change your LR and see what happens? </p>",
      "rawMarkdown": "@ Polaris, I had similar issue as you. I used image size of 64x64 due to memory issues, and I randomly initialized the fully connected layers. My training loss was stunned due to learning rate, not because of my model. Did you try to change your LR and see what happens? "
    },
    {
      "id": 125174,
      "postDate": "2016-06-27T06:33:02.303Z",
      "content": "<p>@Heng CherKeng\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\nWhat's the meaning of &quot;just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.&quot;\nand &quot;As for the last fully connect layers, yo can randomly initialised. &quot;</p>\n\n<p>Since  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning</p>",
      "rawMarkdown": "@Heng CherKeng\r\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\r\nWhat's the meaning of \"just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.\"\r\nand \"As for the last fully connect layers, yo can randomly initialised. \"\r\n\r\nSince  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning\r\n\r\n"
    },
    {
      "id": 125093,
      "postDate": "2016-06-25T18:59:13.587Z",
      "content": "<p>Thanks, which package do you use? Is that caffe?</p>\n\n<p>Thanks,</p>\n\n<p>[quote=Ellen Gao Jian;125004]</p>\n\n<p>Below is my result using vgg16 and batch size 16. I don't think a batch size of 16 would stop convergence.</p>\n\n<p>[quote=Ehsan;124802]</p>\n\n<p>Thanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Thanks, which package do you use? Is that caffe?\r\n\r\nThanks,\r\n\r\n[quote=Ellen Gao Jian;125004]\r\n\r\nBelow is my result using vgg16 and batch size 16. I don't think a batch size of 16 would stop convergence.\r\n\r\n[quote=Ehsan;124802]\r\n\r\nThanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: http://file:///C:/Users/jgao/Desktop/1413561361361513413.JPG\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 125015,
      "postDate": "2016-06-24T16:19:13.280Z",
      "content": "<p>@Jiao Dong:</p>\n\n<p>Like many other before - thank you for sharing your code. Your findings are very helpful.\nLike many others :) I have a question, as well... </p>\n\n<p>How do you use pre-trained models? You are loading the weights from '../input/vgg16_weights.h5' and add an additional softmax layer.</p>\n\n<p>Your last layer has:\n* input-1000 features (last layer of VGG16, the 1000 classes)\n* output-10 features (10 classes our drivers)</p>\n\n<p>When you run model.fit() do you retrain all layers? So you retrain the weights of VGG16 and your additional layer?</p>",
      "rawMarkdown": "@Jiao Dong:\r\n\r\nLike many other before - thank you for sharing your code. Your findings are very helpful.\r\nLike many others :) I have a question, as well... \r\n\r\nHow do you use pre-trained models? You are loading the weights from '../input/vgg16_weights.h5' and add an additional softmax layer.\r\n\r\nYour last layer has:\r\n* input-1000 features (last layer of VGG16, the 1000 classes)\r\n* output-10 features (10 classes our drivers)\r\n\r\nWhen you run model.fit() do you retrain all layers? So you retrain the weights of VGG16 and your additional layer?"
    },
    {
      "id": 124802,
      "postDate": "2016-06-22T13:23:18.550Z",
      "content": "<p>Thanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?</p>",
      "rawMarkdown": "Thanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?"
    },
    {
      "id": 124772,
      "postDate": "2016-06-22T07:14:01.903Z",
      "content": "<p>I believe that you need change the learning rate to a smaller value, such as 1e-4</p>\n\n<p>[quote=JennyYu;124747]</p>\n\n<p>I resized my image to 64 x64, and popped the last layer of hte vgg16, it's not working. My training loss is about 14, accuracy of 10%, so it's doing random guesses for the 10 category. I wonder if the vgg16 model only works well if I main the image size 224 like Jiao Dong just mentioned. Has anyone else got it to work with smaller sized image? </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I believe that you need change the learning rate to a smaller value, such as 1e-4\r\n\r\n[quote=JennyYu;124747]\r\n\r\nI resized my image to 64 x64, and popped the last layer of hte vgg16, it's not working. My training loss is about 14, accuracy of 10%, so it's doing random guesses for the 10 category. I wonder if the vgg16 model only works well if I main the image size 224 like Jiao Dong just mentioned. Has anyone else got it to work with smaller sized image? \r\n\r\n[/quote]\r\n"
    },
    {
      "id": 124747,
      "postDate": "2016-06-21T22:36:46.437Z",
      "content": "<p>I resized my image to 64 x64, and popped the last layer of hte vgg16, it's not working. My training loss is about 14, accuracy of 10%, so it's doing random guesses for the 10 category. I wonder if the vgg16 model only works well if I main the image size 224 like Jiao Dong just mentioned. Has anyone else got it to work with smaller sized image? </p>",
      "rawMarkdown": "I resized my image to 64 x64, and popped the last layer of hte vgg16, it's not working. My training loss is about 14, accuracy of 10%, so it's doing random guesses for the 10 category. I wonder if the vgg16 model only works well if I main the image size 224 like Jiao Dong just mentioned. Has anyone else got it to work with smaller sized image? "
    },
    {
      "id": 124677,
      "postDate": "2016-06-21T06:41:50.927Z",
      "content": "<p>Thanks Jiao Dong. I want to reproduce this result in Tensorflow r8.0 with GPU titanX. But the model always suffer from overfitting and the validation loss can only reach 0.8 without data augmentation and 0.44 with a lot different data augmentation. That's a surprise that you can just use the origin data to get such an amazing result.</p>\n\n<p>I am thinking that is it because of the precision of K40 and TitanX  or the dl tool tensorflow or theano?</p>",
      "rawMarkdown": "Thanks Jiao Dong. I want to reproduce this result in Tensorflow r8.0 with GPU titanX. But the model always suffer from overfitting and the validation loss can only reach 0.8 without data augmentation and 0.44 with a lot different data augmentation. That's a surprise that you can just use the origin data to get such an amazing result.\r\n\r\nI am thinking that is it because of the precision of K40 and TitanX  or the dl tool tensorflow or theano?\r\n"
    },
    {
      "id": 124670,
      "postDate": "2016-06-21T02:20:22.170Z",
      "content": "<p>just change input to 256x256. the convolution filters are now applied to larger area and that's all.\nyou still can use the pretrained vgg16 filters.</p>\n\n<p>As for the last fully connect layers, yo can randomly initialised. I tried 256x256, and now 300x300, 384x384. Will post the results when they are out.</p>\n\n<p>[quote=Ehsan;124631]</p>\n\n<p>How do you train 256x256 with vgg16 pretrained model. The weights was computed for 224x224. Do you only use model without weights?\nThanks</p>\n\n<p>[quote=Heng CherKeng;124465]</p>\n\n<p>i am training a vgg with 16 layers. All layers are fine tuned.\nThe split of the train and test set are based on drivers.\nrandom 4 drivers are used as validation and remaining 22 are used for training.</p>\n\n<p>From various submissions i made, i find that:</p>\n\n<ol>\n<li><p>So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.</p></li>\n<li><p>For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.</p></li>\n</ol>\n\n<p>[quote=Ehsan;124440]</p>\n\n<p>Very interesting!\nDid you keep first 11 layers and train remaining layers?\nDid you split validation based on drivers or just random?</p>\n\n<p>Thank you</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "just change input to 256x256. the convolution filters are now applied to larger area and that's all.\r\nyou still can use the pretrained vgg16 filters.\r\n\r\nAs for the last fully connect layers, yo can randomly initialised. I tried 256x256, and now 300x300, 384x384. Will post the results when they are out.\r\n\r\n[quote=Ehsan;124631]\r\n\r\nHow do you train 256x256 with vgg16 pretrained model. The weights was computed for 224x224. Do you only use model without weights?\r\nThanks\r\n\r\n[quote=Heng CherKeng;124465]\r\n\r\ni am training a vgg with 16 layers. All layers are fine tuned.\r\nThe split of the train and test set are based on drivers.\r\nrandom 4 drivers are used as validation and remaining 22 are used for training.\r\n\r\nFrom various submissions i made, i find that:\r\n\r\n 1. So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.\r\n\r\n 2. For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.\r\n \r\n\r\n[quote=Ehsan;124440]\r\n\r\nVery interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 124669,
      "postDate": "2016-06-21T02:15:23.423Z",
      "content": "<p>You are very welcome ! Glad people succeed to run it :)</p>\n\n<p>This script was my first experience with keras too, thanks for ZFTurbo's original script as well that saved a lot of time. I always believed the best part of building a project is not &quot;take something online and everything just worked out&quot;, it's the days &amp; weeks you spent to debug / customize / hack it. That's how I learn from my projects too :)</p>\n\n<p>[quote=ChrisJung;124594]</p>\n\n<p>Dear Jiao Dong,</p>\n\n<p>Thank you very much for uploading the script.\nI also faced few Memory Errors due to my limited Hardware spec and few syntax errors due to the difference in the library version.\nI resolved those issue by using small batch_size, dropping shuffle by permutation and other minor adjustments for memory efficiency.\nSyntax error was resolved by the post in this forum!</p>\n\n<p>But it was a great debugging experience, which made me learn tons! :)\nIt was my first touch with Keras, and fortunately I succeeded in replicating basic pipeline of your script!\nFew weeks of personal struggle but it was all worth it :)</p>\n\n<p>Thanks again,</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "You are very welcome ! Glad people succeed to run it :)\r\n\r\nThis script was my first experience with keras too, thanks for ZFTurbo's original script as well that saved a lot of time. I always believed the best part of building a project is not \"take something online and everything just worked out\", it's the days & weeks you spent to debug / customize / hack it. That's how I learn from my projects too :)\r\n\r\n[quote=ChrisJung;124594]\r\n\r\nDear Jiao Dong,\r\n\r\nThank you very much for uploading the script.\r\nI also faced few Memory Errors due to my limited Hardware spec and few syntax errors due to the difference in the library version.\r\nI resolved those issue by using small batch_size, dropping shuffle by permutation and other minor adjustments for memory efficiency.\r\nSyntax error was resolved by the post in this forum!\r\n\r\nBut it was a great debugging experience, which made me learn tons! :)\r\nIt was my first touch with Keras, and fortunately I succeeded in replicating basic pipeline of your script!\r\nFew weeks of personal struggle but it was all worth it :)\r\n\r\nThanks again,\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 124646,
      "postDate": "2016-06-20T20:36:10.333Z",
      "content": "<p>@Ehsan\nNo, you resize the images to 224 x 224 or use cropping. Most frameworks have already functions for that.</p>",
      "rawMarkdown": "@Ehsan\r\nNo, you resize the images to 224 x 224 or use cropping. Most frameworks have already functions for that."
    },
    {
      "id": 124631,
      "postDate": "2016-06-20T16:55:25.773Z",
      "content": "<p>How do you train 256x256 with vgg16 pretrained model. The weights was computed for 224x224. Do you only use model without weights?\nThanks</p>\n\n<p>[quote=Heng CherKeng;124465]</p>\n\n<p>i am training a vgg with 16 layers. All layers are fine tuned.\nThe split of the train and test set are based on drivers.\nrandom 4 drivers are used as validation and remaining 22 are used for training.</p>\n\n<p>From various submissions i made, i find that:</p>\n\n<ol>\n<li><p>So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.</p></li>\n<li><p>For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.</p></li>\n</ol>\n\n<p>[quote=Ehsan;124440]</p>\n\n<p>Very interesting!\nDid you keep first 11 layers and train remaining layers?\nDid you split validation based on drivers or just random?</p>\n\n<p>Thank you</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "How do you train 256x256 with vgg16 pretrained model. The weights was computed for 224x224. Do you only use model without weights?\r\nThanks\r\n\r\n[quote=Heng CherKeng;124465]\r\n\r\ni am training a vgg with 16 layers. All layers are fine tuned.\r\nThe split of the train and test set are based on drivers.\r\nrandom 4 drivers are used as validation and remaining 22 are used for training.\r\n\r\nFrom various submissions i made, i find that:\r\n\r\n 1. So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.\r\n\r\n 2. For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.\r\n \r\n\r\n[quote=Ehsan;124440]\r\n\r\nVery interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 124594,
      "postDate": "2016-06-20T06:59:44.280Z",
      "content": "<p>Dear Jiao Dong,</p>\n\n<p>Thank you very much for uploading the script.\nI also faced few Memory Errors due to my limited Hardware spec and few syntax errors due to the difference in the library version.\nI resolved those issue by using small batch_size, dropping shuffle by permutation and other minor adjustments for memory efficiency.\nSyntax error was resolved by the post in this forum!</p>\n\n<p>But it was a great debugging experience, which made me learn tons! :)\nIt was my first touch with Keras, and fortunately I succeeded in replicating basic pipeline of your script!\nFew weeks of personal struggle but it was all worth it :)</p>\n\n<p>Thanks again,</p>",
      "rawMarkdown": "Dear Jiao Dong,\r\n\r\nThank you very much for uploading the script.\r\nI also faced few Memory Errors due to my limited Hardware spec and few syntax errors due to the difference in the library version.\r\nI resolved those issue by using small batch_size, dropping shuffle by permutation and other minor adjustments for memory efficiency.\r\nSyntax error was resolved by the post in this forum!\r\n\r\nBut it was a great debugging experience, which made me learn tons! :)\r\nIt was my first touch with Keras, and fortunately I succeeded in replicating basic pipeline of your script!\r\nFew weeks of personal struggle but it was all worth it :)\r\n\r\nThanks again,"
    },
    {
      "id": 124581,
      "postDate": "2016-06-20T04:39:41.327Z",
      "content": "<p>Hi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:</p>\n\n<p>&quot;Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.&quot;</p>\n\n<p>I've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!</p>\n\n<p>Jenny</p>",
      "rawMarkdown": "Hi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:\r\n\r\n\"Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\"\r\n\r\nI've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!\r\n\r\nJenny"
    },
    {
      "id": 124450,
      "postDate": "2016-06-18T18:06:38.097Z",
      "content": "<p>Did you have a look at the training images?\nMany c9 &quot;talking to passenger&quot; images look like c0 &quot;normal driving&quot; images.\nSo I think it's really hard or near to impossible to get this right for a NN if even a human would classify some images differently.\nThat's also reflected in your confusion matrix.</p>",
      "rawMarkdown": "Did you have a look at the training images?\r\nMany c9 \"talking to passenger\" images look like c0 \"normal driving\" images.\r\nSo I think it's really hard or near to impossible to get this right for a NN if even a human would classify some images differently.\r\nThat's also reflected in your confusion matrix."
    },
    {
      "id": 124440,
      "postDate": "2016-06-18T15:41:50.440Z",
      "content": "<p>Very interesting!\nDid you keep first 11 layers and train remaining layers?\nDid you split validation based on drivers or just random?</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Very interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you"
    },
    {
      "id": 124121,
      "postDate": "2016-06-15T17:22:48.617Z",
      "content": "<p>Thanks ~ so glad to see it worked on other server lol</p>\n\n<p>Yes i used cuda 7.5 and have cudnn installed in server environment. Also since I am using Theano, in configuration file I also enabled fastmath and cnmem.</p>\n\n<p>[quote=scsherm;124118]</p>\n\n<p>The code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Thanks ~ so glad to see it worked on other server lol\r\n\r\nYes i used cuda 7.5 and have cudnn installed in server environment. Also since I am using Theano, in configuration file I also enabled fastmath and cnmem.\r\n\r\n[quote=scsherm;124118]\r\n\r\nThe code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 124092,
      "postDate": "2016-06-15T13:34:59.023Z",
      "content": "<p>So finally is there anyone manage to get the code worked? I have tried but still can't run on my local machine.</p>",
      "rawMarkdown": "So finally is there anyone manage to get the code worked? I have tried but still can't run on my local machine."
    },
    {
      "id": 123700,
      "postDate": "2016-06-13T15:22:25.287Z",
      "content": "<p>You are right, it's the public leaderboard score.</p>\n\n<p>Considering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.</p>\n\n<p>[quote=Duc Nguyen;123697]</p>\n\n<p>Are those the loss scores in your training, or scores in the public leaderboard? \nI guess the later since you had much lower validation loss. \nIn this case, I think I am missing something in my Torch code.\nThanks again.</p>\n\n<p>[quote=Jiao Dong;123690]</p>\n\n<p>At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.</p>\n\n<p>[quote=Duc Nguyen;123685]</p>\n\n<p>@Jiao Dong Thanks for sharing the code.\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "You are right, it's the public leaderboard score.\r\n\r\nConsidering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.\r\n\r\n[quote=Duc Nguyen;123697]\r\n\r\nAre those the loss scores in your training, or scores in the public leaderboard? \r\nI guess the later since you had much lower validation loss. \r\nIn this case, I think I am missing something in my Torch code.\r\nThanks again.\r\n\r\n[quote=Jiao Dong;123690]\r\n\r\nAt 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 123697,
      "postDate": "2016-06-13T15:17:28.953Z",
      "content": "<p>Are those the loss scores in your training, or scores in the public leaderboard? \nI guess the later since you had much lower validation loss. \nIn this case, I think I am missing something in my Torch code.\nThanks again.</p>\n\n<p>[quote=Jiao Dong;123690]</p>\n\n<p>At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.</p>\n\n<p>[quote=Duc Nguyen;123685]</p>\n\n<p>@Jiao Dong Thanks for sharing the code.\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Are those the loss scores in your training, or scores in the public leaderboard? \r\nI guess the later since you had much lower validation loss. \r\nIn this case, I think I am missing something in my Torch code.\r\nThanks again.\r\n\r\n[quote=Jiao Dong;123690]\r\n\r\nAt 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 123691,
      "postDate": "2016-06-13T14:12:46.873Z",
      "content": "<p>good call :)</p>\n\n<p>[quote=Vinh Nguyen;122366]</p>\n\n<p>for the newer version of Keras, you need a slightly different code to do model surgery, i.e. pop out the last layer and insert the new layer</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>[quote=Jiao Dong;121903]</p>\n\n<p>@Manuele Tamburrano</p>\n\n<p>I think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)</p>\n\n<p>Because as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.</p>\n\n<p>I did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "good call :)\r\n\r\n[quote=Vinh Nguyen;122366]\r\n\r\nfor the newer version of Keras, you need a slightly different code to do model surgery, i.e. pop out the last layer and insert the new layer\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\n\r\n\r\n[quote=Jiao Dong;121903]\r\n\r\n@Manuele Tamburrano\r\n\r\nI think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)\r\n\r\nBecause as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.\r\n\r\nI did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(\r\n\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 123690,
      "postDate": "2016-06-13T14:12:29.967Z",
      "content": "<p>At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.</p>\n\n<p>[quote=Duc Nguyen;123685]</p>\n\n<p>@Jiao Dong Thanks for sharing the code.\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 123685,
      "postDate": "2016-06-13T13:23:07.777Z",
      "content": "<p>@Jiao Dong Thanks for sharing the code.\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. </p>",
      "rawMarkdown": "@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n"
    },
    {
      "id": 123554,
      "postDate": "2016-06-12T16:49:06.170Z",
      "content": "<p>Can I confirm if my understanding for the  LB 0.23800 solution is correct or not?</p>\n\n<ul>\n<li><p>No data augmentation was used. The image was just simply resized to 224x224.\n  Color image was used.</p></li>\n<li><p>The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\n  one fold served as validation set and the remaining 7 as training set. The validation set was used \n  to determine when to terminate the training.</p></li>\n</ul>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Can I confirm if my understanding for the  LB 0.23800 solution is correct or not?\r\n\r\n -  No data augmentation was used. The image was just simply resized to 224x224.\r\n      Color image was used.\r\n \r\n - The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\r\n      one fold served as validation set and the remaining 7 as training set. The validation set was used \r\n      to determine when to terminate the training.\r\n\r\nThanks!"
    },
    {
      "id": 122349,
      "postDate": "2016-06-03T08:11:04.823Z",
      "content": "<p>@Jiao Dong Sorry for my previous post, but the fact remains that some people were a little confused. Your code was very helpful for many people:-)</p>",
      "rawMarkdown": "@Jiao Dong Sorry for my previous post, but the fact remains that some people were a little confused. Your code was very helpful for many people:-)"
    },
    {
      "id": 121963,
      "postDate": "2016-05-31T09:32:36.130Z",
      "content": "<p>From my understanding of the data, using validation_split = 0.2 is the right way of solving the problem, which is likely to produce the optimize PB score.</p>\n\n<p>The splitting base on the drivers just give an estimation of how generation the model is, i.e your score should be expected to have the average of the K-fold valid loss. It does not guarantee producing the best models in this case. </p>",
      "rawMarkdown": "From my understanding of the data, using validation_split = 0.2 is the right way of solving the problem, which is likely to produce the optimize PB score.\r\n\r\nThe splitting base on the drivers just give an estimation of how generation the model is, i.e your score should be expected to have the average of the K-fold valid loss. It does not guarantee producing the best models in this case. \r\n"
    },
    {
      "id": 121939,
      "postDate": "2016-05-31T04:52:53.367Z",
      "content": "<p>@Jiao Dong\nOh, I didn't catch that haha. Too bad that my computer is having a memory error with loading the data in RGB, so I can't even tryout your code... :(  Anyway, thx again for the code! Hope you aced your project :) </p>",
      "rawMarkdown": "@Jiao Dong\r\nOh, I didn't catch that haha. Too bad that my computer is having a memory error with loading the data in RGB, so I can't even tryout your code... :(  Anyway, thx again for the code! Hope you aced your project :) "
    },
    {
      "id": 121904,
      "postDate": "2016-05-30T20:47:36.260Z",
      "content": "<p>@Jeong Wook Moon</p>\n\n<p>In the function signature if a parameter is not explicitly passed it would use the default value, like:</p>\n\n<p>def load_train(img_rows, img_cols, color_type=1):</p>\n\n<p>would use grayscale images if i didn't specify to use color images.</p>\n\n<p>At the part right after imports, there's a global variable </p>\n\n<p>color_type_global = 3</p>\n\n<p>which is the parameter actually passed into </p>\n\n<pre><code>train_data, train_target, driver_id, unique_drivers = \\\n    read_and_normalize_and_shuffle_train_data(img_rows, img_cols,\n                                              color_type_global)\n</code></pre>\n\n<p>model = vgg_std16_model(img_rows, img_cols, color_type_global)</p>\n\n<p>test_data, test_id = read_and_normalize_test_data(img_rows, img_cols,\n                                                      color_type_global)</p>\n\n<p>test_data, test_id = read_and_normalize_test_data(img_rows, img_cols,\n                                                      color_type_global)</p>\n\n<p>Did this help for your question ? :)</p>",
      "rawMarkdown": "@Jeong Wook Moon\r\n\r\nIn the function signature if a parameter is not explicitly passed it would use the default value, like:\r\n\r\ndef load_train(img_rows, img_cols, color_type=1):\r\n\r\nwould use grayscale images if i didn't specify to use color images.\r\n\r\nAt the part right after imports, there's a global variable \r\n\r\ncolor_type_global = 3\r\n\r\nwhich is the parameter actually passed into \r\n\r\n    train_data, train_target, driver_id, unique_drivers = \\\r\n        read_and_normalize_and_shuffle_train_data(img_rows, img_cols,\r\n                                                  color_type_global)\r\n\r\nmodel = vgg_std16_model(img_rows, img_cols, color_type_global)\r\n\r\ntest_data, test_id = read_and_normalize_test_data(img_rows, img_cols,\r\n                                                      color_type_global)\r\n\r\ntest_data, test_id = read_and_normalize_test_data(img_rows, img_cols,\r\n                                                      color_type_global)\r\n\r\n\r\nDid this help for your question ? :)"
    },
    {
      "id": 121903,
      "postDate": "2016-05-30T20:31:42.380Z",
      "content": "<p>@Manuele Tamburrano</p>\n\n<p>I think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)</p>\n\n<p>Because as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.</p>\n\n<p>I did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(</p>",
      "rawMarkdown": "@Manuele Tamburrano\r\n\r\nI think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)\r\n\r\nBecause as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.\r\n\r\nI did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(\r\n\r\n"
    },
    {
      "id": 121901,
      "postDate": "2016-05-30T20:21:16.537Z",
      "content": "<p>@rcarson @Manuele Tamburrano</p>\n\n<p>Sorry what i wrote previously is not towards you guys :) , it's the other posts i just saw which seemed to claim I posted the wrong information on purpose to mislead people, especially I didn't think they even tried to setup the environment &amp; code to execute it themselves. </p>\n\n<p>I'm glad you have been trying to execute the code and debug, I guess due to the different versions of libraries and frameworks it might behave a bit differently, that's something I didn't consider when i post it because I was assuming as long as the script is the same it would yield similar result, which seems to be wrong :(</p>\n\n<p>The admin I have been trying to contact is on vacation ... but I would get the server up and running as long as he comes back and see what's the issue.  </p>",
      "rawMarkdown": "@rcarson @Manuele Tamburrano\r\n\r\nSorry what i wrote previously is not towards you guys :) , it's the other posts i just saw which seemed to claim I posted the wrong information on purpose to mislead people, especially I didn't think they even tried to setup the environment & code to execute it themselves. \r\n\r\nI'm glad you have been trying to execute the code and debug, I guess due to the different versions of libraries and frameworks it might behave a bit differently, that's something I didn't consider when i post it because I was assuming as long as the script is the same it would yield similar result, which seems to be wrong :(\r\n\r\nThe admin I have been trying to contact is on vacation ... but I would get the server up and running as long as he comes back and see what's the issue.  \r\n\r\n"
    },
    {
      "id": 121897,
      "postDate": "2016-05-30T19:29:47.417Z",
      "content": "<p>[quote=Jiao Dong;121891]</p>\n\n<p>I have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.</p>\n\n<p>I will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. </p>\n\n<p>I spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.</p>\n\n<p>[/quote]</p>\n\n<p>I never doubted about your code, if you read my comments I wonder if I did something wrong or if my device is different for some reason (float precision) to reproduce your results. I'm sorry about unpleasant comments of users, I can understand they hurt when you share something very useful in a paid competition. Thank you again for your code.</p>\n\n<p>Anyway I spent several hours before giving up to go below 0.4 on LB score, I tried several minor changes on your code, used driver split to detect overfitting, changed parameters, kfold methods, but still stuck on ~0.4.</p>\n\n<p>Now I'm on different road, but I'm still curious on what is the cause to this difference.\n@rcarson, do you use same frameworks and versions of Jiao (anaconda, keras, theano) or the more updated ones? Are you splitting the train by driver or just a random split?\nFrom my tries splitting by drivers results in a better estimate of overfitting, but the LB score is worse</p>",
      "rawMarkdown": "[quote=Jiao Dong;121891]\r\n\r\nI have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.\r\n\r\nI will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. \r\n\r\nI spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.\r\n\r\n\r\n\r\n[/quote]\r\n\r\nI never doubted about your code, if you read my comments I wonder if I did something wrong or if my device is different for some reason (float precision) to reproduce your results. I'm sorry about unpleasant comments of users, I can understand they hurt when you share something very useful in a paid competition. Thank you again for your code.\r\n\r\nAnyway I spent several hours before giving up to go below 0.4 on LB score, I tried several minor changes on your code, used driver split to detect overfitting, changed parameters, kfold methods, but still stuck on ~0.4.\r\n\r\nNow I'm on different road, but I'm still curious on what is the cause to this difference.\r\n@rcarson, do you use same frameworks and versions of Jiao (anaconda, keras, theano) or the more updated ones? Are you splitting the train by driver or just a random split?\r\nFrom my tries splitting by drivers results in a better estimate of overfitting, but the LB score is worse\r\n"
    },
    {
      "id": 120491,
      "postDate": "2016-05-18T16:40:47.303Z",
      "content": "<p>anyone succeded to reproduce that LB score using these code without tweaking parameters?</p>\n\n<p>I did several tries with changes proposed by me and jvipond  and adjusting learning rate but I cannot go beyond ~0.4 on LB score.</p>\n\n<p>I'm moving on different solutions but I'm still wondering why the difference is so big running the same code (probably with newer versions of libraries). So I'm a bit worried that there is something shady in my tries, like float precision or something like that.</p>\n\n<p>So, anyone was able to get better results with the same code?</p>",
      "rawMarkdown": "anyone succeded to reproduce that LB score using these code without tweaking parameters?\r\n\r\nI did several tries with changes proposed by me and jvipond  and adjusting learning rate but I cannot go beyond ~0.4 on LB score.\r\n\r\nI'm moving on different solutions but I'm still wondering why the difference is so big running the same code (probably with newer versions of libraries). So I'm a bit worried that there is something shady in my tries, like float precision or something like that.\r\n\r\nSo, anyone was able to get better results with the same code?"
    },
    {
      "id": 119657,
      "postDate": "2016-05-12T03:52:16.387Z",
      "content": "<p>@ Jiao Dong\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?</p>",
      "rawMarkdown": "@ Jiao Dong\r\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?"
    },
    {
      "id": 119387,
      "postDate": "2016-05-09T18:10:26.933Z",
      "content": "<p>I'm trying a different approach, I edited the load_weights method in engine/topology.py to only load layers originally defined. So I can simply remove the dense layer with 1000 neurons and add a new one with only 10 neurons without having to pop anything.\nIn this way parameters are 40970 that seems correct, the loss is decreasing faster but apparently not as fast as op's loss. I'll let the net to train this night to see where it is going</p>",
      "rawMarkdown": "I'm trying a different approach, I edited the load_weights method in engine/topology.py to only load layers originally defined. So I can simply remove the dense layer with 1000 neurons and add a new one with only 10 neurons without having to pop anything.\r\nIn this way parameters are 40970 that seems correct, the loss is decreasing faster but apparently not as fast as op's loss. I'll let the net to train this night to see where it is going"
    },
    {
      "id": 119384,
      "postDate": "2016-05-09T17:57:33.353Z",
      "content": "<p>I think @jvipond got the solution for the problem, the way you change dense layer is different  in newer versions of Keras. I also spent quite a bit of time messing with libraries on the remote server so my setup might be different as well.</p>\n\n<p>I can take a look of the layers later, but got to finish finals first =__=</p>\n\n<p>[quote=Manuele Tamburrano;119377]</p>\n\n<p>thank you for clarifications, but probably you are remembering something wrong.\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?</p>\n\n<p>And still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons</p>\n\n<p>[quote=Jiao Dong;119371]</p>\n\n<p>I don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here <a href=\"https://anaconda.org/omnia/keras\">https://anaconda.org/omnia/keras</a> (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.</p>\n\n<p>You reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I think @jvipond got the solution for the problem, the way you change dense layer is different  in newer versions of Keras. I also spent quite a bit of time messing with libraries on the remote server so my setup might be different as well.\r\n\r\nI can take a look of the layers later, but got to finish finals first =__=\r\n\r\n[quote=Manuele Tamburrano;119377]\r\n\r\nthank you for clarifications, but probably you are remembering something wrong.\r\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?\r\n\r\nAnd still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\r\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons\r\n\r\n[quote=Jiao Dong;119371]\r\n\r\nI don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here https://anaconda.org/omnia/keras (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.\r\n\r\nYou reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119377,
      "postDate": "2016-05-09T16:43:46.337Z",
      "content": "<p>thank you for clarifications, but probably you are remembering something wrong.\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?</p>\n\n<p>And still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons</p>\n\n<p>[quote=Jiao Dong;119371]</p>\n\n<p>I don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here <a href=\"https://anaconda.org/omnia/keras\">https://anaconda.org/omnia/keras</a> (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.</p>\n\n<p>You reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "thank you for clarifications, but probably you are remembering something wrong.\r\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?\r\n\r\nAnd still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\r\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons\r\n\r\n[quote=Jiao Dong;119371]\r\n\r\nI don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here https://anaconda.org/omnia/keras (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.\r\n\r\nYou reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. \r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119374,
      "postDate": "2016-05-09T16:08:33.080Z",
      "content": "<p>Yep, just need to change prototxt file for the output layer, and add a data layer on top to feed into the model.</p>\n\n<p>We ran out of time to do more experiments before deadline (May 1st) so we didn't try anything further. But seriously, we are using the same GPU but training on his Caffe is like 3~5 times faster than mine</p>\n\n<p>[quote=Bojan Tunguz;119372]</p>\n\n<p>[quote=Jiao Dong;119367]</p>\n\n<p>My teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.</p>\n\n<p>[/quote]</p>\n\n<p>Is that with the same data augmentation and CV as done in this Keras + VGG_16 script?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Yep, just need to change prototxt file for the output layer, and add a data layer on top to feed into the model.\r\n\r\nWe ran out of time to do more experiments before deadline (May 1st) so we didn't try anything further. But seriously, we are using the same GPU but training on his Caffe is like 3~5 times faster than mine\r\n\r\n[quote=Bojan Tunguz;119372]\r\n\r\n[quote=Jiao Dong;119367]\r\n\r\nMy teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.\r\n\r\n\r\n[/quote]\r\n\r\nIs that with the same data augmentation and CV as done in this Keras + VGG_16 script?\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119372,
      "postDate": "2016-05-09T16:03:18.507Z",
      "content": "<p>[quote=Jiao Dong;119367]</p>\n\n<p>My teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.</p>\n\n<p>[/quote]</p>\n\n<p>Is that with the same data augmentation and CV as done in this Keras + VGG_16 script?</p>",
      "rawMarkdown": "[quote=Jiao Dong;119367]\r\n\r\nMy teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.\r\n\r\n\r\n[/quote]\r\n\r\nIs that with the same data augmentation and CV as done in this Keras + VGG_16 script?\r\n"
    },
    {
      "id": 119371,
      "postDate": "2016-05-09T16:01:33.530Z",
      "content": "<p>I don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here <a href=\"https://anaconda.org/omnia/keras\">https://anaconda.org/omnia/keras</a> (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.</p>\n\n<p>You reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. </p>\n\n<p>[quote=Manuele Tamburrano;119370]</p>\n\n<p>yes, I tried to lower learning rate and some small adjustment but the loss keeps decreasing very slowly, I get ~1.3 after 15 epochs.</p>\n\n<p>What version of Keras are you running? The one from pip (it should be 1.0.2) or master compiled from source? If this is the case, what's your last commit hash?</p>\n\n<p>Seems like you are using an old version, because your logs show accuracy values, but you only set &quot;show_accuracy=True&quot; in the fit method, but in the last version you should use metrics=[&quot;accuracy&quot;] in compile method</p>\n\n<p>[quote=Jiao Dong;119366]</p>\n\n<p>There's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.</p>\n\n<p>From replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here https://anaconda.org/omnia/keras (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.\r\n\r\nYou reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. \r\n\r\n[quote=Manuele Tamburrano;119370]\r\n\r\nyes, I tried to lower learning rate and some small adjustment but the loss keeps decreasing very slowly, I get ~1.3 after 15 epochs.\r\n\r\nWhat version of Keras are you running? The one from pip (it should be 1.0.2) or master compiled from source? If this is the case, what's your last commit hash?\r\n\r\nSeems like you are using an old version, because your logs show accuracy values, but you only set \"show_accuracy=True\" in the fit method, but in the last version you should use metrics=[\"accuracy\"] in compile method\r\n\r\n[quote=Jiao Dong;119366]\r\n\r\nThere's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.\r\n\r\nFrom replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119367,
      "postDate": "2016-05-09T15:27:33.340Z",
      "content": "<p>My teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.</p>\n\n<p>[quote=Matteo Presutto;119360]</p>\n\n<p>Did any of you try fine tuning resnet50? The best I can get out of it is 0.85, It starts overfitting after 2-3 iterations</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "My teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.\r\n\r\n[quote=Matteo Presutto;119360]\r\n\r\nDid any of you try fine tuning resnet50? The best I can get out of it is 0.85, It starts overfitting after 2-3 iterations\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119366,
      "postDate": "2016-05-09T15:25:49.143Z",
      "content": "<p>There's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.</p>\n\n<p>From replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?</p>\n\n<p>[quote=Manuele Tamburrano;119348]</p>\n\n<p>Thank you for sharing this code, I was trying to do essentially the same.</p>\n\n<p>Anyway I can't reproduce your results, my loss decreases very slowly and I can't see the fast improvements you talk about in the first ~1000 to ~5000 images in the first epoch.</p>\n\n<p>Did you change anything else but the num_fold and nb_epochs compared to the published code?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "There's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.\r\n\r\nFrom replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?\r\n\r\n[quote=Manuele Tamburrano;119348]\r\n\r\nThank you for sharing this code, I was trying to do essentially the same.\r\n\r\nAnyway I can't reproduce your results, my loss decreases very slowly and I can't see the fast improvements you talk about in the first ~1000 to ~5000 images in the first epoch.\r\n\r\nDid you change anything else but the num_fold and nb_epochs compared to the published code?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119360,
      "postDate": "2016-05-09T14:28:05.787Z",
      "content": "<p>Did any of you try fine tuning resnet50? The best I can get out of it is 0.85, It starts overfitting after 2-3 iterations</p>",
      "rawMarkdown": "Did any of you try fine tuning resnet50? The best I can get out of it is 0.85, It starts overfitting after 2-3 iterations"
    },
    {
      "id": 119348,
      "postDate": "2016-05-09T11:08:20.027Z",
      "content": "<p>Thank you for sharing this code, I was trying to do essentially the same.</p>\n\n<p>Anyway I can't reproduce your results, my loss decreases very slowly and I can't see the fast improvements you talk about in the first ~1000 to ~5000 images in the first epoch.</p>\n\n<p>Did you change anything else but the num_fold and nb_epochs compared to the published code?</p>",
      "rawMarkdown": "Thank you for sharing this code, I was trying to do essentially the same.\r\n\r\nAnyway I can't reproduce your results, my loss decreases very slowly and I can't see the fast improvements you talk about in the first ~1000 to ~5000 images in the first epoch.\r\n\r\nDid you change anything else but the num_fold and nb_epochs compared to the published code?"
    },
    {
      "id": 119186,
      "postDate": "2016-05-07T21:21:34.470Z",
      "content": "<p>LOL it shows how much difference between GPU and CPU in floating point computation  </p>\n\n<p>[quote=Gerard Toonstra;119149]</p>\n\n<p>I posted about how to use a memory mapped file in this thread here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147</a></p>\n\n<p>I also tried to run the network on a CPU, because my GPU doesn't have enough memory. So far, I only have one thread that's working on the data. Estimated time to finish a single epoch for the first fold:</p>\n\n<p>95 days.</p>\n\n<p>LOL! :)</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "LOL it shows how much difference between GPU and CPU in floating point computation  \r\n\r\n[quote=Gerard Toonstra;119149]\r\n\r\nI posted about how to use a memory mapped file in this thread here:\r\n\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147\r\n\r\n\r\nI also tried to run the network on a CPU, because my GPU doesn't have enough memory. So far, I only have one thread that's working on the data. Estimated time to finish a single epoch for the first fold:\r\n\r\n95 days.\r\n\r\nLOL! :)\r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119158,
      "postDate": "2016-05-07T17:47:03.973Z",
      "content": "<p>[quote=Wendy Kan;119155]</p>\n\n<p>Since it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. </p>\n\n<p>[/quote]</p>\n\n<p>Wait a sec, VGG-16 has already been used in several recent Kaggle competitions. Will you retroactively re-evaluate all of those???</p>",
      "rawMarkdown": "[quote=Wendy Kan;119155]\r\n\r\nSince it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. \r\n\r\n  [1]: https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md\r\n\r\n[/quote]\r\n\r\nWait a sec, VGG-16 has already been used in several recent Kaggle competitions. Will you retroactively re-evaluate all of those???"
    },
    {
      "id": 119157,
      "postDate": "2016-05-07T17:44:48.023Z",
      "content": "<p>Oh, that's a pity... :(</p>",
      "rawMarkdown": "Oh, that's a pity... :("
    },
    {
      "id": 119140,
      "postDate": "2016-05-07T15:22:17.943Z",
      "content": "<p>[quote=ZFTurbo;119102]</p>\n\n<p><strong>Abhijay Arora</strong>, If use this code &quot;as is&quot; it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.</p>\n\n<p><a href=\"http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\">http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory</a></p>\n\n<p>True.But with a generator it is possible it seems.Do you have a generator that can be used with this?. I tried with a simple generator. it's taking much time.I'm not sure about it's success.</p>\n\n<p>def generate():</p>\n\n<pre><code>for i in range(2803):  # 2803*8 = 22424  ---&gt; this generator returns as 8 chunks\n    #as i'm trying with gray scale i have it converted to array in memory (My laptop RAM 8GB)\n   #you may need to convert the training set to array &quot;here itself&quot;\n    yield (np.array(X_train[i*8:(i+1)*8],dtype=np.uint8),np.array(Y_newtrain[i*8:(i+1)*8],dtype=np.uint8))\n</code></pre>\n\n<p>for e in range(nb_epoch):\n    print(&quot;epoch %d&quot; % e)\n    for X_batch, Y_batch in generate(): \n        model.fit(X_batch, Y_batch,nb_epoch=1,show_accuracy=True,shuffle=True)</p>\n\n<p>Please help if someone do have a generator.Thanks for the help in advance</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "[quote=ZFTurbo;119102]\r\n\r\n**Abhijay Arora**, If use this code \"as is\" it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.\r\n\r\nhttp://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\r\n\r\nTrue.But with a generator it is possible it seems.Do you have a generator that can be used with this?. I tried with a simple generator. it's taking much time.I'm not sure about it's success.\r\n\r\ndef generate():\r\n\r\n    for i in range(2803):  # 2803*8 = 22424  ---> this generator returns as 8 chunks\r\n        #as i'm trying with gray scale i have it converted to array in memory (My laptop RAM 8GB)\r\n       #you may need to convert the training set to array \"here itself\"\r\n        yield (np.array(X_train[i*8:(i+1)*8],dtype=np.uint8),np.array(Y_newtrain[i*8:(i+1)*8],dtype=np.uint8))\r\n\r\nfor e in range(nb_epoch):\r\n    print(\"epoch %d\" % e)\r\n    for X_batch, Y_batch in generate(): \r\n        model.fit(X_batch, Y_batch,nb_epoch=1,show_accuracy=True,shuffle=True)\r\n\r\nPlease help if someone do have a generator.Thanks for the help in advance\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119107,
      "postDate": "2016-05-07T09:30:33.467Z",
      "content": "<p>@ Abhijay</p>\n\n<p>My immediate guess is that this is caused by the Fully Connected layer (which expects an input at a fixed size) Changing the original input image size will mess up with this layer.</p>",
      "rawMarkdown": "@ Abhijay\r\n\r\nMy immediate guess is that this is caused by the Fully Connected layer (which expects an input at a fixed size) Changing the original input image size will mess up with this layer."
    },
    {
      "id": 119106,
      "postDate": "2016-05-07T09:27:34.157Z",
      "content": "<p>[quote=ZFTurbo;119102]</p>\n\n<p><strong>Abhijay Arora</strong>, If use this code &quot;as is&quot; it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.</p>\n\n<p><a href=\"http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\">http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory</a></p>\n\n<p>[/quote]</p>\n\n<p>Thanks. I tried running the script &quot;as is&quot;, just changed the image size to 64 x 64 . I get the following error while loading the weights:</p>\n\n<pre><code>Exception: Layer weight shape (2048, 4096) not compatible with provided weight shape (25088, 4096)\n</code></pre>\n\n<p>Any idea why?</p>",
      "rawMarkdown": "[quote=ZFTurbo;119102]\r\n\r\n**Abhijay Arora**, If use this code \"as is\" it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.\r\n\r\nhttp://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\r\n\r\n[/quote]\r\n\r\nThanks. I tried running the script \"as is\", just changed the image size to 64 x 64 . I get the following error while loading the weights:\r\n\r\n    Exception: Layer weight shape (2048, 4096) not compatible with provided weight shape (25088, 4096)\r\n\r\nAny idea why?\r\n"
    },
    {
      "id": 119104,
      "postDate": "2016-05-07T08:14:10.590Z",
      "content": "<p>It is interesting that you managed to achieve such a good score without overfitting to the persons in the training set without having any holdout validation set with persons unseen during training. Was it just alot of trial and error of parameters? </p>",
      "rawMarkdown": "It is interesting that you managed to achieve such a good score without overfitting to the persons in the training set without having any holdout validation set with persons unseen during training. Was it just alot of trial and error of parameters? "
    },
    {
      "id": 119071,
      "postDate": "2016-05-07T00:56:22.170Z",
      "content": "<p>@Jiao</p>\n\n<p>Just wanted to quickly thank you for posting the code and for such a very thorough and detailed writeup. I have a few general questions, but I'll leave those for later after I've played with your code a bit. </p>",
      "rawMarkdown": "@Jiao\r\n\r\nJust wanted to quickly thank you for posting the code and for such a very thorough and detailed writeup. I have a few general questions, but I'll leave those for later after I've played with your code a bit. "
    },
    {
      "id": 119062,
      "postDate": "2016-05-06T23:46:15.517Z",
      "content": "<p>Hi Jiao Dong,</p>\n\n<p>Can you share the parameters of you model like vgg16? Without gpu, it will take several days for 20 epoch. Or just share the visualization of first layer. Very curious about how the first layer will look like. </p>\n\n<p>BTW, is this a bug in? I see &quot;train_drivers&quot; and &quot;test_drivers&quot; never be used in the loop:</p>\n\n<p>for train_drivers, test_drivers in kf:\n        num_fold += 1\n        print('Start KFold number {} from {}'.format(num_fold, nfolds))\n        # print('Split train: ', len(X_train), len(Y_train))\n        # print('Split valid: ', len(X_valid), len(Y_valid))\n        # print('Train drivers: ', unique_list_train)\n        # print('Test drivers: ', unique_list_valid)\n        # model = create_model_v1(img_rows, img_cols, color_type_global)\n        # model = vgg_bn_model(img_rows, img_cols, color_type_global)\n        model = vgg_std16_model(img_rows, img_cols, color_type_global)</p>\n\n<pre><code>    model.fit(train_data, train_target, batch_size=batch_size,\n              nb_epoch=nb_epoch,\n              show_accuracy=True, verbose=1,\n              validation_split=split, shuffle=True)\n</code></pre>",
      "rawMarkdown": "Hi Jiao Dong,\r\n\r\nCan you share the parameters of you model like vgg16? Without gpu, it will take several days for 20 epoch. Or just share the visualization of first layer. Very curious about how the first layer will look like. \r\n\r\nBTW, is this a bug in? I see \"train_drivers\" and \"test_drivers\" never be used in the loop:\r\n\r\nfor train_drivers, test_drivers in kf:\r\n        num_fold += 1\r\n        print('Start KFold number {} from {}'.format(num_fold, nfolds))\r\n        # print('Split train: ', len(X_train), len(Y_train))\r\n        # print('Split valid: ', len(X_valid), len(Y_valid))\r\n        # print('Train drivers: ', unique_list_train)\r\n        # print('Test drivers: ', unique_list_valid)\r\n        # model = create_model_v1(img_rows, img_cols, color_type_global)\r\n        # model = vgg_bn_model(img_rows, img_cols, color_type_global)\r\n        model = vgg_std16_model(img_rows, img_cols, color_type_global)\r\n\r\n        model.fit(train_data, train_target, batch_size=batch_size,\r\n                  nb_epoch=nb_epoch,\r\n                  show_accuracy=True, verbose=1,\r\n                  validation_split=split, shuffle=True)\r\n\r\n\r\n"
    },
    {
      "id": 119009,
      "postDate": "2016-05-06T16:52:04.073Z",
      "content": "<p>That's interesting to see, unfortunately I don't know the solution either. This script is one of the many I've tried that converged well, if you discovered any insight of this model / script, please keep me updated :) </p>\n\n<p>[quote=threecourse;118999]</p>\n\n<p>@rcarson</p>\n\n<p>In my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \nVal_loss is not reliable due to split by image not driver.</p>\n\n<p>In addition, by printing model.summary(), number of weights of last layer is 10010.\nIt indicates layers are not connected properly. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "That's interesting to see, unfortunately I don't know the solution either. This script is one of the many I've tried that converged well, if you discovered any insight of this model / script, please keep me updated :) \r\n\r\n[quote=threecourse;118999]\r\n\r\n@rcarson\r\n\r\nIn my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \r\nVal_loss is not reliable due to split by image not driver.\r\n\r\nIn addition, by printing model.summary(), number of weights of last layer is 10010.\r\nIt indicates layers are not connected properly. \r\n\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119007,
      "postDate": "2016-05-06T16:37:09.520Z",
      "content": "<p>Yep, I didn't use driver information in training and shuffled my training data on purpose. </p>\n\n<p>However I didn't run any model based on driver information with high epoch.</p>\n\n<p>[quote=ZFTurbo;118979]</p>\n\n<p>I think on these lines you actually lost the connection between drivers and images:</p>\n\n<pre><code>perm = permutation(len(train_target))\ntrain_data = train_data[perm]\ntrain_target = train_target[perm]\n</code></pre>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Yep, I didn't use driver information in training and shuffled my training data on purpose. \r\n\r\nHowever I didn't run any model based on driver information with high epoch.\r\n\r\n[quote=ZFTurbo;118979]\r\n\r\nI think on these lines you actually lost the connection between drivers and images:\r\n\r\n    perm = permutation(len(train_target))\r\n    train_data = train_data[perm]\r\n    train_target = train_target[perm]\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 119000,
      "postDate": "2016-05-06T16:16:07.823Z",
      "content": "<p>Thank you. I'll let you know if I run into the same thing or not</p>",
      "rawMarkdown": "Thank you. I'll let you know if I run into the same thing or not"
    },
    {
      "id": 118994,
      "postDate": "2016-05-06T15:43:18.947Z",
      "content": "<p>@threecourse, hi, i'm running it now, what do you mean by 'doesn't work'? CV score is too bad?</p>\n\n<p>Best,</p>",
      "rawMarkdown": "@threecourse, hi, i'm running it now, what do you mean by 'doesn't work'? CV score is too bad?\r\n\r\nBest,"
    },
    {
      "id": 118985,
      "postDate": "2016-05-06T14:58:09.030Z",
      "content": "<p>It didn't work for me, anyone success?</p>\n\n<p>I think in the script CV is split by image, not by driver.</p>",
      "rawMarkdown": "It didn't work for me, anyone success?\r\n\r\nI think in the script CV is split by image, not by driver."
    },
    {
      "id": 118952,
      "postDate": "2016-05-06T10:07:48.163Z",
      "content": "<p>@saihttam</p>\n\n<p>Try reducing the batch size</p>",
      "rawMarkdown": "@saihttam\r\n\r\nTry reducing the batch size"
    },
    {
      "id": 118920,
      "postDate": "2016-05-06T05:04:49.833Z",
      "content": "<p>[quote=Jiao Dong;118880]...it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU[/quote]</p>\n\n<p>Is this with TensorFlow 0.8? There are supposed to be speed improvements in the latest version.</p>\n\n<p>Many thanks for posting your script, looking forward to trying it out!</p>",
      "rawMarkdown": "[quote=Jiao Dong;118880]...it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU[/quote]\r\n\r\nIs this with TensorFlow 0.8? There are supposed to be speed improvements in the latest version.\r\n\r\nMany thanks for posting your script, looking forward to trying it out!\r\n"
    },
    {
      "id": 118917,
      "postDate": "2016-05-06T03:53:31.607Z",
      "content": "<p>@Jiao, thanks for your sharing. This is really awesome!</p>",
      "rawMarkdown": "@Jiao, thanks for your sharing. This is really awesome!"
    },
    {
      "id": 118946,
      "postDate": "2016-05-06T09:07:49.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2008185,
      "postDate": "2022-10-28T20:08:13.300Z",
      "content": "<p>thanks for explaning  this simple VGG16</p>",
      "rawMarkdown": "thanks for explaning  this simple VGG16"
    },
    {
      "id": 125776,
      "postDate": "2016-07-02T13:05:29.413Z",
      "content": "<p>Thanks for sharing&#65281;</p>",
      "rawMarkdown": "Thanks for sharing！"
    },
    {
      "id": 124829,
      "postDate": "2016-06-22T16:38:40.640Z",
      "content": "<p>Thanks for the suggestion, Yu Hai</p>",
      "rawMarkdown": "Thanks for the suggestion, Yu Hai"
    },
    {
      "id": 124675,
      "postDate": "2016-06-21T04:05:24.963Z",
      "content": "<p>Thank you, Jiao Dong</p>",
      "rawMarkdown": "Thank you, Jiao Dong"
    }
  ],
  "comments": [
    {
      "id": 119383,
      "author_name": "jvipond",
      "author_url": "",
      "post_date": "2016-05-09T17:48:11.463000",
      "content": "<p>[quote=Manuele Tamburrano;119377]</p>\n\n<p>thank you for clarifications, but probably you are remembering something wrong.\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?</p>\n\n<p>And still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons</p>\n\n<p>[/quote]</p>\n\n<p>That's the problem. It seems to be an issue with newer versions of Keras. See here <a href=\"https://github.com/fchollet/keras/issues/2371\">https://github.com/fchollet/keras/issues/2371</a>.\nI experienced the same thing with the script converging very slowly to a much higher than reported loss and having incorrect param numbers and then I changed the pop layer code to:</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>as suggested in the issue and now it seems to be converging a lot faster and with a lower loss.</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 121891,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-05-30T19:06:58.657000",
      "content": "<p>I have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.</p>\n\n<p>I will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. </p>\n\n<p>I spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 118886,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-05-05T22:52:00.203000",
      "content": "<p>In addition, I came across this online course while searching for information about CNN. I have been following it for a while, it's  the best introductory course of CNN online.</p>\n\n<p><a href=\"http://cs231n.stanford.edu/\">http://cs231n.stanford.edu/</a></p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 121127,
      "author_name": "Zhao W",
      "author_url": "",
      "post_date": "2016-05-24T07:23:05.467000",
      "content": "<p>I ran the script and got ~0.6LB with single model.</p>\n\n<p>Here is my guess why the author can get 0.2LB.</p>\n\n<p>According to the author, </p>\n\n<ul>\n<li>&quot;I have attached the original python code that got 0.23800 loss score, using 8 folds with 15 epochs.&quot;</li>\n<li>&quot;If you want to run 'fast' experiments, I got a LB 0.32640 with only 2 folds, 3 epoch each&quot;</li>\n</ul>\n\n<p>In main.py line 365, he splits drivers into train_drivers and test_drivers with KFold. However, when training the model in line 378, he uses the full training set. So it seems to me that the cross-validation part actually just trains the model on same data many times with different init. </p>\n\n<p>Maybe training more models with different init can improve LB score.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 119189,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-05-07T21:39:26.627000",
      "content": "<p>Good point, the file I used do have restricted use, so for people who is considering to participate this competition seriously you should be aware of licence issue. </p>\n\n<p>But there're still plenty of models released based on ImageNet that are unrestricted, like <a href=\"http://caffe.berkeleyvision.org/model_zoo.html#bvlc-model-license\">Caffe's Model Zoo BVLC Model license</a>  and <a href=\"https://github.com/KaimingHe/deep-residual-networks\">ResNet's Third-Party re-implementations (including Kaggle)</a> , if you insist using VGG-16 in keras, I just found admin's updated response <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread/119178#post119178\">Here at post #28</a></p>\n\n<p>I have already moved on to other projects since end of April, but at least I showed what loss score you can get with a particular pre-trained model. There are many other models you can try. From admin's reminder of copyright, please make sure you chose the right model without commercial restriction. GL, HF :)</p>\n\n<p>[quote=Wendy Kan;119155]</p>\n\n<p>Hi all, </p>\n\n<p>Someone in the community flagged the usage of <a href=\"https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md\">VGG-16</a> here. We looked into the license and found this in their disclaimer:</p>\n\n<blockquote>\n  <p>license: <a href=\"http://creativecommons.org/licenses/by-nc/4.0/\">http://creativecommons.org/licenses/by-nc/4.0/</a>\n  (non-commercial use only)</p>\n</blockquote>\n\n<p>Since it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. </p>\n\n<p>[/quote]</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 125242,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-06-28T02:09:37.653000",
      "content": "<p>there are a few ways to do.  It is a &quot;trial and error&quot; and see which would work best.\nGiven a pretrained CNN network, we wish to change the size of the filter and weights.</p>\n\n<p>E.g. for making them larger,  we can:</p>\n\n<ol>\n<li>pad new values with zeros   </li>\n<li><p>fill new values with random values. The magnitude of the random values are important.\n You need to ensure:</p>\n\n<ul><li><p>distribution (old_output) = distribution(new_output) </p></li>\n<li><p>where: old_output = function(old_weight, old_input) ,  new_output =function(new_weight, old_input), and  function = conv or inner product</p></li></ul></li>\n</ol>\n\n<p>Assuming Gaussian, you just need to ensure mean and std of old_output and new_output are the same. Hence the new  random values can be just random Gaussian noise of appropriate std.</p>\n\n<p>for 224x224 to 256x256, I think you can just pad with zero and try first. I works for me as well.</p>\n\n<p>For reference, refer to:</p>\n\n<ul>\n<li><a href=\"http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization\">http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization</a></li>\n<li><a href=\"http://arxiv.org/abs/1511.06422\">http://arxiv.org/abs/1511.06422</a></li>\n<li><a href=\"http://deepdish.io/2015/02/24/network-initialization/\">http://deepdish.io/2015/02/24/network-initialization/</a></li>\n</ul>\n\n<p>the key of training deep networks is to make sure signals can propagate forward and backward.\nThe weights values (and data values) cannot be too large or too small. If too large, it will grow infinitely large and leads to explosion (you will see #NAN in training loss). If too small, it will grow infinitely small, aka the problem of diminishing gradient.</p>\n\n<p>[quote=Polaris;125174]</p>\n\n<p>@Heng CherKeng\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\nWhat's the meaning of &quot;just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.&quot;\nand &quot;As for the last fully connect layers, yo can randomly initialised. &quot;</p>\n\n<p>Since  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning</p>\n\n<p>[/quote]</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 119077,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-05-07T02:00:10.940000",
      "content": "<p>When you train, could you try comment out the initializing with the pretrain vgg16.pkl and just use random initialization? by doing so, we know how much it really helps with the pretrained net.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 118999,
      "author_name": "threecourse",
      "author_url": "",
      "post_date": "2016-05-06T16:14:20.677000",
      "content": "<p>@rcarson</p>\n\n<p>In my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \nVal_loss is not reliable due to split by image not driver.</p>\n\n<p>In addition, by printing model.summary(), number of weights of last layer is 10010.\nIt indicates layers are not connected properly. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 119102,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2016-05-07T08:02:25.173000",
      "content": "<p><strong>Abhijay Arora</strong>, If use this code &quot;as is&quot; it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.</p>\n\n<p><a href=\"http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\">http://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 129608,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2016-08-01T00:02:03.060000",
      "content": "<p>[quote=Mike Kim;129605]</p>\n\n<p>The geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 \n[/quote]</p>\n\n<p>It should not be an issue, at least because with softmax output 0 prediction can not happen due to the fact that </p>\n\n<blockquote>\n  <p>exp[x] = 0 &lt;=&gt; x = - Infinity</p>\n</blockquote>\n\n<p>And -Infinity should not appear because </p>\n\n<blockquote>\n  <p>min(float number stored in computer) &gt; -Infinity</p>\n</blockquote>\n\n<p>Although I can imaging zero probability appearing due to some rounding.</p>\n\n<p>Just looked through some submissions for this competition, I see very small numbers, but I do not see any zeros.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129347,
      "author_name": "Andrey Vykhodtsev",
      "author_url": "",
      "post_date": "2016-07-29T00:18:40.353000",
      "content": "<p>As I found out theano on GPU would only accelerate operations with float32  as stated here :\n<a href=\"http://deeplearning.net/software/theano/tutorial/using_gpu.html\">http://deeplearning.net/software/theano/tutorial/using_gpu.html</a></p>\n\n<blockquote>\n  <p>What Can Be Accelerated on the GPU\n  The performance characteristics will change as we continue to optimize our implementations, and vary from device to device, but to give a rough idea of what to expect right now:</p>\n  \n  <p>Only computations with float32 data-type can be accelerated. Better support for float64 is expected in upcoming hardware but float64 computations are still relatively slow (Jan 2010).</p>\n</blockquote>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 129101,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2016-07-26T17:45:08.410000",
      "content": "<p>@tetemin:</p>\n\n<ol>\n<li>Adam definitely helps. </li>\n<li>TensorFlow is roughly twice slower than  Theano =&gt; if you change your backend at aws it will spin faster. </li>\n<li>You  are finetuning =&gt; learning rate should be really small. I use 1e-5, 1e-6.</li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 128799,
      "author_name": "NelsonChen",
      "author_url": "",
      "post_date": "2016-07-24T03:28:34.617000",
      "content": "<p>Hm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 128790,
      "author_name": "Ferris Wu",
      "author_url": "",
      "post_date": "2016-07-24T00:20:48.670000",
      "content": "<p>@NelsonChen</p>\n\n<p>With SWAP partition, you can run 224x224 on a 16GB memory PC, with just a bit change on test data loading process. </p>\n\n<p>Each epoch is taking 900s under GTX 1070 &amp; CUDA 8.0, 4x faster than AWS g2 instance.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 126416,
      "author_name": "PengPai",
      "author_url": "",
      "post_date": "2016-07-08T09:05:36.113000",
      "content": "<p>@VZ\nIt is simple to resolve your problem. Please put 'train' and 'test' inside a new fold named 'imgs'. Done.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 126154,
      "author_name": "VZ",
      "author_url": "",
      "post_date": "2016-07-06T21:11:35.650000",
      "content": "<p>I just tried for the first time your script, but I am getting this error on:\ntrain_data = train_data.transpose((0, 3, 1, 2))</p>\n\n<p>ValueError: axes don't match array</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125403,
      "author_name": "Cogitae _ Thomas Soumarmon",
      "author_url": "",
      "post_date": "2016-06-29T07:02:07.567000",
      "content": "<p>Thanks everyone for sharing knowledge and know-how.</p>\n\n<p>Hello @Jadiel, I had a quick look at your script and you should be able to improve the top_model by using the weights for the 2 Dense 4096 layers from the VGG16 weigths and also use some init like &quot;he_normal&quot; for the last Dense 10 layer</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125330,
      "author_name": "Ehsan M. Ardehaly",
      "author_url": "",
      "post_date": "2016-06-28T17:25:38.167000",
      "content": "<p>You can use numpy.memmap and map a file to memory. It is pretty fast, and does not harm your training performance.</p>\n\n<p>[quote=JennyYu;125250]</p>\n\n<p>@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. </p>\n\n<p>I learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  </p>\n\n<p>Good luck to you.</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 124668,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-06-21T02:07:25.677000",
      "content": "<p>From my own understanding of normalization, it applies when you have multiple features in numerical format, sometimes in different numerical range. You would like to know for a particular data point, how much feature X deviates from other data points for the same attribute X. Sometimes the numerical value of feature would mislead you, for example, without any normalization a feature value 200 might appear to be twice as significant as feature value 100 of another data point, but if your average value for that feature field is 10,000, they might in fact be nearly the same, equally trivial. So for color images people often divide each pixel channel by 255 (0xFF) to normalize each color channel, also make sure no color channel's value would dominate others. </p>\n\n<p>For vgg pre-trained models, it's a bit different. If you refer to the <a href=\"https://arxiv.org/pdf/1409.1556.pdf\"> Original VGG Net paper by Karen Simonyan &amp; Andrew Zisserman</a> , in section 2.1 ARCHITECTURE, &quot;During training, the input to our ConvNets is a fixed-size 224 &#215; 224 RGB image. The only preprocessing we do is subtracting the mean RGB value, computed on the training set, from each pixel&quot;, they trained vgg model from 14+ millions of images on <a href=\"http://www.image-net.org/\">ImageNet</a>, and those &quot;magic numbers&quot; are the mean pixel value for each channel based on images they used for training, serving as benchmarks for each channel.</p>\n\n<p>[quote=JennyYu;124581]</p>\n\n<p>Hi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:</p>\n\n<p>&quot;Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.&quot;</p>\n\n<p>I've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!</p>\n\n<p>Jenny</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 124465,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-06-18T22:14:38.650000",
      "content": "<p>i am training a vgg with 16 layers. All layers are fine tuned.\nThe split of the train and test set are based on drivers.\nrandom 4 drivers are used as validation and remaining 22 are used for training.</p>\n\n<p>From various submissions i made, i find that:</p>\n\n<ol>\n<li><p>So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.</p></li>\n<li><p>For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.</p></li>\n</ol>\n\n<p>[quote=Ehsan;124440]</p>\n\n<p>Very interesting!\nDid you keep first 11 layers and train remaining layers?\nDid you split validation based on drivers or just random?</p>\n\n<p>Thank you</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 124118,
      "author_name": "scsherm",
      "author_url": "",
      "post_date": "2016-06-15T17:13:35.170000",
      "content": "<p>The code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?</p>\n\n<pre><code>model.layers.pop()\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 123902,
      "author_name": "Duc Nguyen",
      "author_url": "",
      "post_date": "2016-06-14T12:12:48.167000",
      "content": "<p>I see. Its nice that you have very good scores even without data augmentation. I got some improvements with my Torch implementation but I guess I'm still missing some important parts. </p>\n\n<p>Btw, the main reason for having higher log loss in the public leaderboard than validation loss is not the difference in the numbers of image but is the fact that you are not splitting train/val images based on drivers. I once did the same thing at the beginning and also got very low validation loss. Now I'm splitting train/val set based on drivers and the validation loss seems to be a good indicator of the public leaderboard score.</p>\n\n<p>[quote=Jiao Dong;123700]</p>\n\n<p>You are right, it's the public leaderboard score.</p>\n\n<p>Considering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.</p>\n\n<p>[quote=Duc Nguyen;123697]</p>\n\n<p>Are those the loss scores in your training, or scores in the public leaderboard? \nI guess the later since you had much lower validation loss. \nIn this case, I think I am missing something in my Torch code.\nThanks again.</p>\n\n<p>[quote=Jiao Dong;123690]</p>\n\n<p>At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.</p>\n\n<p>[quote=Duc Nguyen;123685]</p>\n\n<p>@Jiao Dong Thanks for sharing the code.\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. </p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 123590,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-06-12T22:59:57.133000",
      "content": "<ul>\n<li>No data augmentation was used. The image was just simply resized to 224x224. Color image was used.</li>\n</ul>\n\n<p>Yes, exactly.</p>\n\n<ul>\n<li>The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\n  one fold served as validation set and the remaining 7 as training set. The validation set was used \n  to determine when to terminate the training.</li>\n</ul>\n\n<p>The final model is the average of 8 models. The training phase is a bit different from your description: for each fold, you went over your entire training dataset and generate a model file, then there's a random split of training dataset such that random 85% is used for training and 15% is for validating, between each fold, it's highly likely that they are using different subset of training data / validation data, with different ordering.</p>\n\n<p>[quote=Heng CherKeng;123554]</p>\n\n<p>Can I confirm if my understanding for the  LB 0.23800 solution is correct or not?</p>\n\n<ul>\n<li><p>No data augmentation was used. The image was just simply resized to 224x224.\n  Color image was used.</p></li>\n<li><p>The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\n  one fold served as validation set and the remaining 7 as training set. The validation set was used \n  to determine when to terminate the training.</p></li>\n</ul>\n\n<p>Thanks!</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 121907,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-05-30T20:53:15.863000",
      "content": "<p>I tried vgg-19 with very limited amount of time spent on it ( I remember it was like 2 days before deadline so I can't possibly finish a good run) and its learning convergence gave me nearly the same validation loss with 15 epochs. I would say it might not be much better, but not worse than vgg-16 as well.</p>\n\n<p>[quote=kuan chen;119657]</p>\n\n<p>@ Jiao Dong\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 121862,
      "author_name": "Jeong Wook Moon",
      "author_url": "",
      "post_date": "2016-05-30T12:47:34.587000",
      "content": "<p>@Jiao Dong\nThx for sharing your code on the forum! But I have a question about your code.\nYou said that you used colored 224x224 images, but in your code, the variable color_type = 1 in your model. And when I try to load the pre-trained VGG weights and retrain, the model sends out an error message stating that the shape of the model weights are not compatible, (64,1,3,3) with (64,3,3,3).\nIt would be really great if you can help me out here. Thx!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 121137,
      "author_name": "Keiku",
      "author_url": "",
      "post_date": "2016-05-24T08:19:57.597000",
      "content": "<p>I don't use this script. I voted down because it seemed that this script didn't produce ~0.2 LB score. This will cause an confusion. The correct information should be shared.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 121065,
      "author_name": "René Scheibe",
      "author_url": "",
      "post_date": "2016-05-23T12:05:19.553000",
      "content": "<p>[quote=anokas;119526]</p>\n\n<p>I still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?</p>\n\n<p>My epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here</p>\n\n<p>[/quote]</p>\n\n<p>With a Titan X + CUDNN 7.5 (Winograd convolution algorithm) + Theano 0.8.2 (see .theanorc Settings below) I can get to 550s per epoch.</p>\n\n<pre><code>[dnn.conv]                                       \nalgo_fwd = time_once\nalgo_bwd_data = time_once\nalgo_bwd_filter = time_once\n</code></pre>\n\n<p>Removing the ZeroPadding2D layer and using Convolution2D with border_mode='same' instead of 'valid' reduces the time per epoch to 430s.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119526,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2016-05-11T07:13:52.643000",
      "content": "<p>I still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?</p>\n\n<p>My epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119476,
      "author_name": "Manuele Tamburrano",
      "author_url": "",
      "post_date": "2016-05-10T15:32:35.980000",
      "content": "<p>My run with 25 epochs and 3 KFolds with the fix above ended with these values of loss and accuracy:</p>\n\n<p>loss: 0.0017 - acc: 0.9994 - val_loss: 0.0168 - val_acc: 0.9949</p>\n\n<p>Final LB score is ~0.6</p>\n\n<p>Anyway the validation is splitted on images and not by drivers so the validation loss is inaccurate, I'll run the train again with data splitted by driver ids to check when overfit occurs</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119370,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T15:50:10.673000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119159,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T17:48:38.157000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119119,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T12:30:28.413000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 118954,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T10:31:07.840000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 118923,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T05:48:21.343000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125289,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T13:00:23.817000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 125004,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-24T13:44:47.490000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 124422,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-18T10:11:15.093000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 122366,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-03T12:27:57.367000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 121894,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-30T19:15:23.057000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 119155,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T17:40:25.703000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 119149,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T16:40:17.600000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 118979,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T14:22:23.693000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 120764,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-20T12:27:02.723000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118924,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T05:51:29.750000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119097,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T07:04:57.090000",
      "content": "",
      "votes": -3,
      "replies": []
    },
    {
      "id": 129638,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-08-01T07:03:04.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129615,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-08-01T01:41:36.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129605,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-31T23:12:41.033000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129560,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-31T07:57:55.760000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129555,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-31T03:40:45.687000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129551,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-31T01:58:36.900000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129488,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-30T10:28:07.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129470,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-30T03:44:56.773000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129448,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-29T19:31:15.973000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129326,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-28T19:53:29.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129193,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-27T12:41:18.263000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129136,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-27T00:16:21.507000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129099,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-26T17:17:21.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129095,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-26T17:11:01.657000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129093,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-26T16:54:07.450000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129028,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-26T00:45:47.093000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128803,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-24T05:42:55.540000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128787,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-23T22:33:15.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128665,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-22T12:46:29.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128631,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-22T04:25:37.097000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128616,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-21T23:16:03.507000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128600,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-21T17:32:09.697000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128547,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-21T04:52:32.037000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128470,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-20T11:02:22.573000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128461,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-20T10:10:32.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128437,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-20T06:37:51.260000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128222,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-18T15:33:51.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 127357,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-14T01:01:51.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 127053,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-13T06:29:03.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126452,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-08T14:48:28.643000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126450,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-08T14:44:36.500000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126374,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-08T00:08:30.477000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126369,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-07T23:22:56.317000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126334,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-07T19:13:43.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126115,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-06T11:33:40.007000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126113,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-06T11:17:23.960000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126097,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-06T05:30:49.987000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126093,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-06T04:39:35.130000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125843,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-03T11:24:01.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125690,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T15:52:34.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125688,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T15:45:29.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125680,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T15:05:47.510000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125679,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T15:05:23.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125678,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T15:01:07.247000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125674,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T14:46:49.177000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125488,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-29T22:22:15.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125427,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-29T11:28:42.707000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125362,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T22:45:42.090000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125250,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T04:11:05.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125243,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T02:27:08.750000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125241,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T02:05:40.477000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125240,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T01:54:45.253000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125238,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-28T01:32:35.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125207,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-27T16:48:42.250000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125202,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-27T15:28:02.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125174,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-27T06:33:02.303000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125093,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-25T18:59:13.587000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125015,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-24T16:19:13.280000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124802,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-22T13:23:18.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124772,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-22T07:14:01.903000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124747,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-21T22:36:46.437000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124677,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-21T06:41:50.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124670,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-21T02:20:22.170000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124669,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-21T02:15:23.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124646,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-20T20:36:10.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124631,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-20T16:55:25.773000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124594,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-20T06:59:44.280000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124581,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-20T04:39:41.327000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124450,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-18T18:06:38.097000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124440,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-18T15:41:50.440000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124121,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-15T17:22:48.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124092,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-15T13:34:59.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 123700,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-13T15:22:25.287000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 123697,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-13T15:17:28.953000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 123691,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-13T14:12:46.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 123690,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-13T14:12:29.967000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 123685,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-13T13:23:07.777000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 123554,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-12T16:49:06.170000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 122349,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-03T08:11:04.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121963,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-31T09:32:36.130000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121939,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-31T04:52:53.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121904,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-30T20:47:36.260000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121903,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-30T20:31:42.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121901,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-30T20:21:16.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121897,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-30T19:29:47.417000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120491,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-18T16:40:47.303000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119657,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-12T03:52:16.387000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119387,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T18:10:26.933000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119384,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T17:57:33.353000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119377,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T16:43:46.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119374,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T16:08:33.080000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119372,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T16:03:18.507000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119371,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T16:01:33.530000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119367,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T15:27:33.340000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119366,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T15:25:49.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119360,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T14:28:05.787000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119348,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-09T11:08:20.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119186,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T21:21:34.470000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119158,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T17:47:03.973000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119157,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T17:44:48.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119140,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T15:22:17.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119107,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T09:30:33.467000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119106,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T09:27:34.157000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119104,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T08:14:10.590000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119071,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-07T00:56:22.170000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119062,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T23:46:15.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119009,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T16:52:04.073000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119007,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T16:37:09.520000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119000,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T16:16:07.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118994,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T15:43:18.947000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118985,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T14:58:09.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118952,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T10:07:48.163000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118920,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T05:04:49.833000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118917,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T03:53:31.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 118946,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T09:07:49.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2008185,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-28T20:08:13.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125776,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-02T13:05:29.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124829,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-22T16:38:40.640000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 124675,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-06-21T04:05:24.963000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118880": "I have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)\r\n\r\nI call it \"simple solution\" since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.\r\n\r\nFeel free to ask if you have any questions !\r\n\r\nThe code is based on ZFTurbo's thread\r\n[Keras Sample Code][1]\r\n\r\n\r\nFor a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. \r\n\r\nAll my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)\r\n\r\nSome experience I gained from my experiments:\r\n\r\n - Pre-trained model\r\n    - In Keras, there are many pre-trained models available online, like [VGG-16][2] and [VGG-19][3]. In Caffe there're also  [ResNet-50,101,152 by Kaiming He, MSRA][4]. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.\r\n \r\n    - Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\r\n    - When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.\r\n   - From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization [Caffe][5] would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)\r\n   - Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training \"quick and dirty\" models probably will not work very well.\r\n\r\n - During Training\r\n\r\n   - Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease **VERY QUICKLY** within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.\r\n\r\n   - Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see **a trend of decreasing validation loss with minor fluctuation**. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \r\n![# of epochs is too small][6]\r\n![Much better # of epochs, keep an eye on the trend of validation loss][7]\r\n![Trend][8]\r\n\r\n - Some other small changes I made\r\n\r\n   - Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.\r\n\r\n   - I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.\r\n   - Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.\r\n   - I tried using [darknet][9] to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like \"if you see that person in back seat, then.....\" However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\r\n  [2]: https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n  [3]: https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\r\n  [4]: https://github.com/KaimingHe/deep-residual-networks\r\n  [5]: http://caffe.berkeleyvision.org/\r\n  [6]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\r\n  [7]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\r\n  [8]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\r\n  [9]: http://pjreddie.com/darknet/yolo/",
    "119383": "[quote=Manuele Tamburrano;119377]\r\n\r\nthank you for clarifications, but probably you are remembering something wrong.\r\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?\r\n\r\nAnd still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\r\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons\r\n\r\n[/quote]\r\n\r\nThat's the problem. It seems to be an issue with newer versions of Keras. See here https://github.com/fchollet/keras/issues/2371.\r\nI experienced the same thing with the script converging very slowly to a much higher than reported loss and having incorrect param numbers and then I changed the pop layer code to:\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\nas suggested in the issue and now it seems to be converging a lot faster and with a lower loss.",
    "121891": "I have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.\r\n\r\nI will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. \r\n\r\nI spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.\r\n\r\n",
    "118886": "In addition, I came across this online course while searching for information about CNN. I have been following it for a while, it's  the best introductory course of CNN online.\r\n\r\n[http://cs231n.stanford.edu/][1]\r\n\r\n\r\n  [1]: http://cs231n.stanford.edu/",
    "121127": "I ran the script and got ~0.6LB with single model.\r\n\r\nHere is my guess why the author can get 0.2LB.\r\n\r\nAccording to the author, \r\n\r\n - \"I have attached the original python code that got 0.23800 loss score, using 8 folds with 15 epochs.\"\r\n - \"If you want to run 'fast' experiments, I got a LB 0.32640 with only 2 folds, 3 epoch each\"\r\n\r\nIn main.py line 365, he splits drivers into train_drivers and test_drivers with KFold. However, when training the model in line 378, he uses the full training set. So it seems to me that the cross-validation part actually just trains the model on same data many times with different init. \r\n\r\nMaybe training more models with different init can improve LB score.",
    "119189": "Good point, the file I used do have restricted use, so for people who is considering to participate this competition seriously you should be aware of licence issue. \r\n\r\nBut there're still plenty of models released based on ImageNet that are unrestricted, like [Caffe's Model Zoo BVLC Model license][1]  and [ResNet's Third-Party re-implementations (including Kaggle)][2] , if you insist using VGG-16 in keras, I just found admin's updated response [Here at post #28][3]\r\n\r\nI have already moved on to other projects since end of April, but at least I showed what loss score you can get with a particular pre-trained model. There are many other models you can try. From admin's reminder of copyright, please make sure you chose the right model without commercial restriction. GL, HF :)\r\n\r\n[quote=Wendy Kan;119155]\r\n\r\nHi all, \r\n\r\nSomeone in the community flagged the usage of [VGG-16][4] here. We looked into the license and found this in their disclaimer:\r\n\r\n> license: http://creativecommons.org/licenses/by-nc/4.0/\r\n> (non-commercial use only)\r\n\r\nSince it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. \r\n\r\n[/quote]\r\n\r\n\r\n  [1]: http://caffe.berkeleyvision.org/model_zoo.html#bvlc-model-license\r\n  [2]: https://github.com/KaimingHe/deep-residual-networks\r\n  [3]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20141/official-pre-trained-models-and-external-data-thread/119178#post119178\r\n  [4]: https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md",
    "125242": "there are a few ways to do.  It is a \"trial and error\" and see which would work best.\r\nGiven a pretrained CNN network, we wish to change the size of the filter and weights.\r\n\r\nE.g. for making them larger,  we can:\r\n\r\n 1. pad new values with zeros   \r\n 2. fill new values with random values. The magnitude of the random values are important.\r\n     You need to ensure:\r\n\r\n - distribution (old_output) = distribution(new_output) \r\n   \r\n\r\n - where: old_output = function(old_weight, old_input) ,  new_output =function(new_weight, old_input), and  function = conv or inner product\r\n\r\nAssuming Gaussian, you just need to ensure mean and std of old_output and new_output are the same. Hence the new  random values can be just random Gaussian noise of appropriate std.\r\n\r\nfor 224x224 to 256x256, I think you can just pad with zero and try first. I works for me as well.\r\n\r\nFor reference, refer to:\r\n\r\n - http://andyljones.tumblr.com/post/110998971763/an-explanation-of-xavier-initialization\r\n - http://arxiv.org/abs/1511.06422\r\n - http://deepdish.io/2015/02/24/network-initialization/\r\n\r\nthe key of training deep networks is to make sure signals can propagate forward and backward.\r\nThe weights values (and data values) cannot be too large or too small. If too large, it will grow infinitely large and leads to explosion (you will see #NAN in training loss). If too small, it will grow infinitely small, aka the problem of diminishing gradient.\r\n\r\n\r\n[quote=Polaris;125174]\r\n\r\n@Heng CherKeng\r\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\r\nWhat's the meaning of \"just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.\"\r\nand \"As for the last fully connect layers, yo can randomly initialised. \"\r\n\r\nSince  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning\r\n\r\n\r\n\r\n[/quote]\r\n",
    "119077": "When you train, could you try comment out the initializing with the pretrain vgg16.pkl and just use random initialization? by doing so, we know how much it really helps with the pretrained net.",
    "118999": "@rcarson\r\n\r\nIn my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \r\nVal_loss is not reliable due to split by image not driver.\r\n\r\nIn addition, by printing model.summary(), number of weights of last layer is 10010.\r\nIt indicates layers are not connected properly. \r\n",
    "119102": "**Abhijay Arora**, If use this code \"as is\" it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.\r\n\r\nhttp://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\r\n\r\n",
    "129608": "[quote=Mike Kim;129605]\r\n\r\nThe geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 \r\n[/quote]\r\n\r\nIt should not be an issue, at least because with softmax output 0 prediction can not happen due to the fact that \r\n\r\n> exp[x] = 0 <=> x = - Infinity\r\n\r\n\r\nAnd -Infinity should not appear because \r\n\r\n> min(float number stored in computer) > -Infinity\r\n\r\nAlthough I can imaging zero probability appearing due to some rounding.\r\n\r\nJust looked through some submissions for this competition, I see very small numbers, but I do not see any zeros.",
    "129347": "As I found out theano on GPU would only accelerate operations with float32  as stated here :\r\nhttp://deeplearning.net/software/theano/tutorial/using_gpu.html\r\n\r\n> What Can Be Accelerated on the GPU\r\nThe performance characteristics will change as we continue to optimize our implementations, and vary from device to device, but to give a rough idea of what to expect right now:\r\n\r\n>Only computations with float32 data-type can be accelerated. Better support for float64 is expected in upcoming hardware but float64 computations are still relatively slow (Jan 2010).",
    "129101": "@tetemin:\r\n\r\n 1. Adam definitely helps. \r\n 2. TensorFlow is roughly twice slower than  Theano => if you change your backend at aws it will spin faster. \r\n 3. You  are finetuning => learning rate should be really small. I use 1e-5, 1e-6.",
    "128799": "Hm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?",
    "128790": "@NelsonChen\r\n\r\nWith SWAP partition, you can run 224x224 on a 16GB memory PC, with just a bit change on test data loading process. \r\n\r\nEach epoch is taking 900s under GTX 1070 & CUDA 8.0, 4x faster than AWS g2 instance.",
    "126416": "@VZ\r\nIt is simple to resolve your problem. Please put 'train' and 'test' inside a new fold named 'imgs'. Done.",
    "126154": "I just tried for the first time your script, but I am getting this error on:\r\ntrain_data = train_data.transpose((0, 3, 1, 2))\r\n\r\nValueError: axes don't match array",
    "125403": "Thanks everyone for sharing knowledge and know-how.\r\n\r\n\r\nHello @Jadiel, I had a quick look at your script and you should be able to improve the top_model by using the weights for the 2 Dense 4096 layers from the VGG16 weigths and also use some init like \"he_normal\" for the last Dense 10 layer",
    "125330": "You can use numpy.memmap and map a file to memory. It is pretty fast, and does not harm your training performance.\r\n\r\n[quote=JennyYu;125250]\r\n\r\n@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. \r\n\r\nI learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  \r\n\r\nGood luck to you.\r\n\r\n[/quote]\r\n",
    "124668": "From my own understanding of normalization, it applies when you have multiple features in numerical format, sometimes in different numerical range. You would like to know for a particular data point, how much feature X deviates from other data points for the same attribute X. Sometimes the numerical value of feature would mislead you, for example, without any normalization a feature value 200 might appear to be twice as significant as feature value 100 of another data point, but if your average value for that feature field is 10,000, they might in fact be nearly the same, equally trivial. So for color images people often divide each pixel channel by 255 (0xFF) to normalize each color channel, also make sure no color channel's value would dominate others. \r\n\r\nFor vgg pre-trained models, it's a bit different. If you refer to the [ Original VGG Net paper by Karen Simonyan & Andrew Zisserman][1] , in section 2.1 ARCHITECTURE, \"During training, the input to our ConvNets is a fixed-size 224 × 224 RGB image. The only preprocessing we do is subtracting the mean RGB value, computed on the training set, from each pixel\", they trained vgg model from 14+ millions of images on [ImageNet][2], and those \"magic numbers\" are the mean pixel value for each channel based on images they used for training, serving as benchmarks for each channel.\r\n\r\n\r\n[quote=JennyYu;124581]\r\n\r\nHi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:\r\n\r\n\"Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\"\r\n\r\nI've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!\r\n\r\nJenny\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: https://arxiv.org/pdf/1409.1556.pdf\r\n  [2]: http://www.image-net.org/",
    "124465": "i am training a vgg with 16 layers. All layers are fine tuned.\r\nThe split of the train and test set are based on drivers.\r\nrandom 4 drivers are used as validation and remaining 22 are used for training.\r\n\r\nFrom various submissions i made, i find that:\r\n\r\n 1. So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.\r\n\r\n 2. For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.\r\n \r\n\r\n[quote=Ehsan;124440]\r\n\r\nVery interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you\r\n\r\n[/quote]\r\n",
    "124118": "The code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))",
    "123902": "I see. Its nice that you have very good scores even without data augmentation. I got some improvements with my Torch implementation but I guess I'm still missing some important parts. \r\n\r\n\r\nBtw, the main reason for having higher log loss in the public leaderboard than validation loss is not the difference in the numbers of image but is the fact that you are not splitting train/val images based on drivers. I once did the same thing at the beginning and also got very low validation loss. Now I'm splitting train/val set based on drivers and the validation loss seems to be a good indicator of the public leaderboard score.\r\n\r\n[quote=Jiao Dong;123700]\r\n\r\nYou are right, it's the public leaderboard score.\r\n\r\nConsidering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.\r\n\r\n[quote=Duc Nguyen;123697]\r\n\r\nAre those the loss scores in your training, or scores in the public leaderboard? \r\nI guess the later since you had much lower validation loss. \r\nIn this case, I think I am missing something in my Torch code.\r\nThanks again.\r\n\r\n[quote=Jiao Dong;123690]\r\n\r\nAt 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "123590": "- No data augmentation was used. The image was just simply resized to 224x224. Color image was used.\r\n\r\nYes, exactly.\r\n\r\n - The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\r\n      one fold served as validation set and the remaining 7 as training set. The validation set was used \r\n      to determine when to terminate the training.\r\n\r\nThe final model is the average of 8 models. The training phase is a bit different from your description: for each fold, you went over your entire training dataset and generate a model file, then there's a random split of training dataset such that random 85% is used for training and 15% is for validating, between each fold, it's highly likely that they are using different subset of training data / validation data, with different ordering.\r\n\r\n[quote=Heng CherKeng;123554]\r\n\r\nCan I confirm if my understanding for the  LB 0.23800 solution is correct or not?\r\n\r\n -  No data augmentation was used. The image was just simply resized to 224x224.\r\n      Color image was used.\r\n \r\n - The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\r\n      one fold served as validation set and the remaining 7 as training set. The validation set was used \r\n      to determine when to terminate the training.\r\n\r\nThanks!\r\n\r\n[/quote]\r\n",
    "121907": "I tried vgg-19 with very limited amount of time spent on it ( I remember it was like 2 days before deadline so I can't possibly finish a good run) and its learning convergence gave me nearly the same validation loss with 15 epochs. I would say it might not be much better, but not worse than vgg-16 as well.\r\n\r\n[quote=kuan chen;119657]\r\n\r\n@ Jiao Dong\r\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?\r\n\r\n[/quote]\r\n",
    "121862": "@Jiao Dong\r\nThx for sharing your code on the forum! But I have a question about your code.\r\nYou said that you used colored 224x224 images, but in your code, the variable color_type = 1 in your model. And when I try to load the pre-trained VGG weights and retrain, the model sends out an error message stating that the shape of the model weights are not compatible, (64,1,3,3) with (64,3,3,3).\r\nIt would be really great if you can help me out here. Thx!\r\n\r\n",
    "121137": "I don't use this script. I voted down because it seemed that this script didn't produce ~0.2 LB score. This will cause an confusion. The correct information should be shared.",
    "121065": "[quote=anokas;119526]\r\n\r\nI still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?\r\n\r\nMy epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here\r\n\r\n[/quote]\r\n\r\nWith a Titan X + CUDNN 7.5 (Winograd convolution algorithm) + Theano 0.8.2 (see .theanorc Settings below) I can get to 550s per epoch.\r\n\r\n    [dnn.conv]                                       \r\n    algo_fwd = time_once\r\n    algo_bwd_data = time_once\r\n    algo_bwd_filter = time_once\r\n\r\nRemoving the ZeroPadding2D layer and using Convolution2D with border_mode='same' instead of 'valid' reduces the time per epoch to 430s.",
    "119526": "I still have ~2 loss after training 1 epoch on my machine, while the screenshots show just 0.5 loss. Were the screenshots taken when learning rate was set to 0.1 or am I doing something wrong?\r\n\r\nMy epochs are also taking just 800 seconds on a TITAN X, not 1700 like in OP. Not sure if I missed something here",
    "119476": "My run with 25 epochs and 3 KFolds with the fix above ended with these values of loss and accuracy:\r\n\r\nloss: 0.0017 - acc: 0.9994 - val_loss: 0.0168 - val_acc: 0.9949\r\n\r\nFinal LB score is ~0.6\r\n\r\nAnyway the validation is splitted on images and not by drivers so the validation loss is inaccurate, I'll run the train again with data splitted by driver ids to check when overfit occurs",
    "119370": "yes, I tried to lower learning rate and some small adjustment but the loss keeps decreasing very slowly, I get ~1.3 after 15 epochs.\r\n\r\nWhat version of Keras are you running? The one from pip (it should be 1.0.2) or master compiled from source? If this is the case, what's your last commit hash?\r\n\r\nSeems like you are using an old version, because your logs show accuracy values, but you only set \"show_accuracy=True\" in the fit method, but in the last version you should use metrics=[\"accuracy\"] in compile method\r\n\r\n[quote=Jiao Dong;119366]\r\n\r\nThere's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.\r\n\r\nFrom replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?\r\n\r\n\r\n[/quote]\r\n",
    "119159": "@Wendy, but in [here][1], it says `License: unrestricted use`. Does it mean that it is restricted in Caffe, but unrestricted in Keras?\r\n\r\nForgive me if I asked dumb questions. If VGG-16 is not allowed, does it mean that we are not allowed to use the pretrained weights during training, but we can still use its architecture?  Or both are not allowed?\r\n\r\n  [1]: https://github.com/albertomontesg/keras-model-zoo/tree/master/models/VGG-16",
    "119119": "@Abhijay, if you change the image size, you can try the code I posted [here][1]. Add the following code before fullly connected layer. \r\n\r\n    assert os.path.exists(weights_path), 'Model weights not found (see \"weights_path\" variable in script).'\r\n    f = h5py.File(weights_path)\r\n    for k in range(f.attrs['nb_layers']):\r\n    if k >= len(model.layers):\r\n        # we don't look at the last (fully-connected) layers in the savefile\r\n        break\r\n    g = f['layer_{}'.format(k)]\r\n    weights = [g['param_{}'.format(p)] for p in range(g.attrs['nb_params'])]\r\n    model.layers[k].set_weights(weights)\r\n    f.close()\r\n    print('Model loaded.')\r\n\r\nHope it helps. \r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20466/vgg-16-keras/117026#post117026",
    "118954": "is there a smaller pretrained models?\r\n\r\nlooks like it's too big for my hardware too.",
    "118923": "\r\nI don't remember =.=  \r\n\r\nMy roommate had been using tensorflow recently as well with similar problem. From papers and documents he read, tensorflow originally works upon google's infrastructure therefore for the open-source version a lot of optimizations and features are still not implemented, but I would expect its performance issue to be resolved in later version. News I read last week also mentioned google's DeepMind officially announced to switch to tensorflow from Torch for their projects, I would consider the open-source version of tensorflow is still under active development and optimization. \r\n\r\nAnyways, this script has left plenty of things not yet implemented , like image localization (properly deal with the other person at back seat), changing CNN structure, locking weights, changing activation function, changing learning parameter for particular layer, etc. From what I read from authors of VGG and ResNet, it's also worth trying to combine multiple models to improve overall classification performance. Even with the same single vgg-16 model, there are a whole bunch of things you can do other than just taking average. \r\n\r\nWish this script is helpful to be a good starting point for later experiments based on pre-trained models :)\r\n[quote=gauss256;118920]\r\n\r\n[quote=Jiao Dong;118880]...it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU[/quote]\r\n\r\nIs this with TensorFlow 0.8? There are supposed to be speed improvements in the latest version.\r\n\r\nMany thanks for posting your script, looking forward to trying it out!\r\n\r\n\r\n[/quote]\r\n",
    "125289": "the attachment shows the results for 224x224 and 256x256 and the ensemble of the two.\r\n\r\n\r\n[quote=Vladimir Iglovikov;125238]\r\n\r\nAfter some tweaks and modifications I got 0.19 at public LB.\r\n\r\nDid anyone try to use bigger image size and not 224x224 ?\r\n\r\n[/quote]\r\n",
    "125004": "Below is my result using vgg16 and batch size 16. I don't think a batch size of 16 would stop convergence.\r\n\r\n[quote=Ehsan;124802]\r\n\r\nThanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: http://file:///C:/Users/jgao/Desktop/1413561361361513413.JPG",
    "124422": "@ Jiao Dong Thank for your code and method. I followed your idea and did the implementation in C/C++ myself. Here is my implementation:\r\n\r\n 1.  Use all train samples to finetune a VGG16\r\n 2.  Repeat for eight times:\r\n    - Finetune (1) using a subset of train samples. Use the remaining samples as validation to decide when   to stop training\r\n 3. Submission results is the average of eight results from (2)\r\n\r\nFinal results on leader board is 0.27857.\r\n\r\nHere are some differences of my implementation compared to @ Jiao Dong's\r\n\r\n -  use data argumentation (scale shift and rotate)\r\n -  very short training iterations. Fine tunning in step (2) above is limited to 2 epoch.\r\n\r\n\r\nHere is break down of performances of each models and their combination and the confusion matrix of the model later. \r\n\r\nI note that leader board results is very dependent of the selection of the validation drivers and how to terminate the training. (I think this is due insufficient training data provided). With some tweaks, you can improve the leader board score to about 0.23 as reported by @ Jiao Dong.\r\n\r\nTo go beyond 0.23, what i did is to increase input image size to 256x256. you can still use the vgg16 imageNet pretrained model for initialization. Finally I ensembled the results of 224x224 and 256x256. This gives about 0.20.\r\n\r\n\r\n\r\n",
    "122366": "for the newer version of Keras, you need a slightly different code to do model surgery, i.e. pop out the last layer and insert the new layer\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\n\r\n\r\n[quote=Jiao Dong;121903]\r\n\r\n@Manuele Tamburrano\r\n\r\nI think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)\r\n\r\nBecause as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.\r\n\r\nI did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(\r\n\r\n\r\n\r\n[/quote]\r\n",
    "121894": "Thank you very much for sharing your code @Jiao Dong, you don't have to record the video. Sorry for posting this late, but we want to confirm with the exact same code we got 0.21 LB. It does have some variance from device to device and from run to run. But as long as it converges, the result is very stable.",
    "119155": "Hi all, \r\n\r\nSomeone in the community flagged the usage of [VGG-16][1] here. We looked into the license and found this in their disclaimer:\r\n\r\n> license: http://creativecommons.org/licenses/by-nc/4.0/\r\n> (non-commercial use only)\r\n\r\nSince it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. \r\n\r\n  [1]: https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md",
    "119149": "I posted about how to use a memory mapped file in this thread here:\r\n\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147\r\n\r\n\r\nI also tried to run the network on a CPU, because my GPU doesn't have enough memory. So far, I only have one thread that's working on the data. Estimated time to finish a single epoch for the first fold:\r\n\r\n95 days.\r\n\r\nLOL! :)\r\n",
    "118979": "I think on these lines you actually lost the connection between drivers and images:\r\n\r\n    perm = permutation(len(train_target))\r\n    train_data = train_data[perm]\r\n    train_target = train_target[perm]",
    "120764": "I find the proposed solution and corresponding LB score suspicious. With a pretrained AlexNet or pretrained VGG you will not get below 0.4 LB score. I have a hard time seeing what was done differently from just a common VGG here to get near 0.2. Note that the leap between 0.4LB and 0.2LB is very high.",
    "118924": "I'm impressed you pulled off this score without knowing about Karpathy's course.",
    "119097": "[quote=Jiao Dong;118880]\r\n\r\nI have attached the original python code that got 0.23800 loss score,  using 8 folds with 15 epochs. (sorry im not quite sure how to upload scripts in kaggle)\r\n\r\nI call it \"simple solution\" since there's few things I tried that actually worked well, and in this script there're only a few changes made base on the original. So, there's still big room for improvement.\r\n\r\nFeel free to ask if you have any questions !\r\n\r\nThe code is based on ZFTurbo's thread\r\n[Keras Sample Code][1]\r\n\r\n\r\nFor a school machine learning project with limited time for implementation, starter scripts and discussions in this forum saved tremendous amount of time for me to experiment more models / ideas, so thank you, and here's my own two cents for the community. \r\n\r\nAll my experiments were executed on my school's server with Tesla K40c GPU, average time of training and testing as I can recall...... ~1760s for training per epoch, ~2200s for testing each model. So the loss 0.23800 script with 8 folds and 15 epochs took approximately 8 * 15 * 1760 + 8 * 2200 (secs) ~=  63.5 (hrs)\r\n\r\nSome experience I gained from my experiments:\r\n\r\n - Pre-trained model\r\n    - In Keras, there are many pre-trained models available online, like [VGG-16][2] and [VGG-19][3]. In Caffe there're also  [ResNet-50,101,152 by Kaiming He, MSRA][4]. For ResNet I haven't found pre-trained models directly compatible with Keras, also keep in mind I have read about posts oberserving a loss of accuracy if you convert a caffe model file to keras.\r\n \r\n    - Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\r\n    - When load a pre-trained model, my advice is to keep the original input image and layer structure exactly the same, so in my script it uses colored 224x224 image. You can then manipulate layers as you want, like a straight-forward way of using pre-trained model for this problem is very simple, setup your model with exactly the same structure, load weights, then pop the last later of Dense(1000) since we only need to classify 10 classes, and add a Dense(10) layer to it.\r\n   - From my experience messing with pre-trained models, I recommend using Keras for fast-prototyping to test your ideas; however for fine-tuning and customization [Caffe][5] would be a better choice, it is highly popular in academic research, most ImageNet models uses Caffe with their pre-trained model released to public, functionalities like setting up layer-specific features as well as training time. (My VGG_16 took about 3 days, my teammates ResNet-50 on Caffe finished execution overnight)\r\n   - Fine-tuning pre-trained model usually would take hours, days, even weeks, if you want your model to converge to reasonable loss for submission, training \"quick and dirty\" models probably will not work very well.\r\n\r\n - During Training\r\n\r\n   - Learning rate is critical for convergence, since it is your step size during forward-backward propagation in neural network. The first experiment of mine using pre-trained model I set my initial learning rate to 0.1, the next morning when it finished executing,  final loss on LB is 21+....... A rule of thumb, keep looking at the first ~1000 to ~5000 images in your first epoch. You should have a high training loss (~4 to 5) in the beginning with validation accuracy of ~0.1, but it should decrease **VERY QUICKLY** within the first couple hundred pictures, otherwise it is pretty much pointless to keep running your model and you should change your learning rate. In my script I tried couple times and found 0.001 works pretty well, but you can definitely find better and more accurate initial learning rates.\r\n\r\n   - Since there's noise in training data, we should choose the right epochs that converge to a good model without overfitting. I found that during cross validation, keep an eye on your validation loss of between each epoch gave you valuable information about your training process. If you have a good number of epochs and learning rate, as your training goes to deeper epochs you should see **a trend of decreasing validation loss with minor fluctuation**. Refer to the pictures I have attached to see what you should expect with different epochs. The best validation loss I got is consistently less than 0.01, but since it's near deadline I didn't try to go any deeper or train with higher number of folds. \r\n![# of epochs is too small][6]\r\n![Much better # of epochs, keep an eye on the trend of validation loss][7]\r\n![Trend][8]\r\n\r\n - Some other small changes I made\r\n\r\n   - Before loading training data into memory I generated a permutation that shuffles the order of training images / labels that preserves their 1-1 mapping relation, and I shuffle training data between each epoch to keep it as random as possible. I didn't experiment extensively about the idea of cross validation based on drivers so I can't say how it would work, but I think it makes more sense since in testing you are only given a picture without knowing who that person is; thus in training phase you should avoid fitting your model and do cross validation aware of particular driver as well. There might be a smart way to make use of driver ids given in training data, but I haven't figured out or tried yet.\r\n\r\n   - I use a bigger batch size whenever possible. The server I used ran out of memory when I tried 128 so I settled with 64. But as much as I know about batch normalization, having larger batch size in training makes more sense to me.\r\n   - Keras works with Theano and TensorFlow backend. The server I used have two Tesla K40c GPU but by default Theano would only use one of them, the TensorFlow automatically use both. After spending couple hours dealing with the zero padding bug in TensorFlow , I successfully changed the backend and observed it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU ..... a sad story ....... Later I figured in case of multiple folds, you can let each GPU ran a process that trains same model but saves the model file with different names, and later run the test_and_submit with all the model files you generated to make use of multiple GPUs.\r\n   - I tried using [darknet][9] to perform localization to crop the region that only includes driver. (You got to change their C-library a little bit to save the cropped image instead of just drawing squares on them) The initiative is for a lot of pictures you can see another person in the back-seat and I am afraid it might mislead our model a little bit, like \"if you see that person in back seat, then.....\" However with very primitive implementation of cropping and do training / testing with cropped image, it did not work very well. Mostly because after looking through our cropped images many of them became off-centered and few images were even cropped terribly with only part of driver's body, thus introduced more variance in our data. But I still think it's an interesting idea, you got to ensure the quality of localized / cropped images.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras\r\n  [2]: https://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n  [3]: https://gist.github.com/baraldilorenzo/8d096f48a1be4a2d660d\r\n  [4]: https://github.com/KaimingHe/deep-residual-networks\r\n  [5]: http://caffe.berkeleyvision.org/\r\n  [6]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4196/3.JPG\r\n  [7]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4197/6.JPG\r\n  [8]: https://www.kaggle.com/blobs/download/forum-message-attachment-files/4199/NumberofEpoch.png\r\n  [9]: http://pjreddie.com/darknet/yolo/\r\n\r\n[/quote]\r\n\r\n\r\n\r\nFirstly, thank you for sharing your ideas!\r\nI tried running your script as is, but got a memory error. I think this is because the image size is too large(224 x224). I have been training on sizes 64 x 64. Any tips on how can I get 224 x 224 sized images to fit into my RAM (4 GB)?  What changes need to be made to existing code?\r\nThanks!",
    "129638": "Probably geometry mean should be done with following procedure:\r\n1) fix 0 -> 0.00001 and 1 -> 0.99999\r\n2) Apply geometry mean on new fixed data.",
    "129615": "It seems you're correct regarding the Keras models. I checked the minimums and it looks something like: 2.393357e-35 rather than 0. I originally thought it was 0 because R's summary function rounds on the display by default.",
    "129605": "The geometric mean is worse than mean because any row (test observation) with a single model's 0 probability prediction goes to 0 automatically regardless of the other models (e.g. given 3 models, the geomean of 0.99,1,0 is still 0). Kaggle bounds logloss, but there's still a big penalty for guessing 0 and being wrong. ",
    "129560": "**HeshamEraqi**, in all my experiments geom was worse than mean.",
    "129555": "[quote=HeshamEraqi;129551]\r\n\r\nDid anyone test`merge_several_folds_mean` versus `merge_several_folds_geom` on LB ?\r\nDo they give same LB as expected ?\r\n\r\n\r\n[/quote]\r\nI tried geom and get match worse score, you also have to renormalize after that as it no longer sums to 1 per row. \r\n\r\nI also found if you use softmax with temperature >1 you can get better result on LB, but this seemt ot work only for non ensembled model. ",
    "129551": "Did anyone test`merge_several_folds_mean` versus `merge_several_folds_geom` on LB ?\r\nDo they give same LB as expected ?\r\n",
    "129488": "Getting 0.325 with mean of 13 folds. I attach code mainly taken from here and there.",
    "129470": "Getting 0.36797 on LB with first fold and 0.23746 with mean of all 8 fold.  Do you have similar results? Usually ensemble helps just a bit...",
    "129448": "[quote=Vladimir Iglovikov;129101]\r\n\r\n@tetemin:\r\n\r\n 1. Adam definitely helps. \r\n 2. TensorFlow is roughly twice slower than  Theano => if you change your backend at aws it will spin faster. \r\n 3. You  are finetuning => learning rate should be really small. I use 1e-5, 1e-6.\r\n\r\n[/quote]\r\n\r\n\r\nYes, thats interesting, that most papers suggest using SGD for fine-tuning, but Adam and Nadam(Adam version with netsterov momentum) works match better in my cases it gets to 95% in 3 iterations compared to 6 iterations of SGD.  Deep learning is the new field and some recommendations out there become obsolete once some one tires not to use them. \r\n\r\nI guess it only works with pretreated bottleneck classifier, haven't tried without.\r\n",
    "129326": "Can somebody explain what advantage of using float32/float64 instead of uint8? For example here:\r\n\r\n    train_target = np_utils.to_categorical(train_target, 10)\r\n    train_data = train_data.astype('float32')",
    "129193": "[quote=Vladimir Iglovikov;129101]\r\n\r\n@tetemin:\r\n\r\n 1. Adam definitely helps. \r\n 2. TensorFlow is roughly twice slower than  Theano => if you change your backend at aws it will spin faster. \r\n 3. You  are finetuning => learning rate should be really small. I use 1e-5, 1e-6.\r\n\r\n[/quote]\r\n\r\nThanks Vladimir, trying with Adam now. I managed to get Tensorflow up to Theano speed by transposing the weights and using tf dimension ordering instead of th.\r\n\r\nAny other tips on how you managed to get this down to the 0.2-0.3 range of scores since I don't have that much time to experiment. I'm considering trying the following:\r\n\r\n- Training the last fully connected layers with bottleneck results first for the classification problem first and then fine-tuning the whole network. Is that worth trying or did you get good results just from end-to-end fine-tuning from the beginning?\r\n- Augmenting data by cropping out random patches, is this necessary to get a decent score?\r\n",
    "129136": "[quote=NelsonChen;128799]\r\n\r\nHm it seems that I am also having the memory issue when we declare the train_data array be of type float32. Did anyone just leave the type as uint8, would this cause problems for the model?\r\n\r\n[/quote]\r\n\r\nI was stuck with exactly the same problems but have solved them by following the advice of @Ferris, ie changing the SWAP partition size.\r\n\r\nI found the following link http://askubuntu.com/questions/178712/how-to-increase-swap-space?noredirect=1&lq=1 helpful.\r\n\r\nI was getting out of memory errors with 16Gb of ram and an a 5GB swap file. Increasing the swap file to 32GB removes the memory problems. In order to resize my swap partition I had to boot from a recovery disk, resize the partition above the existing swap partition and then expand the swap.\r\n\r\nIf, like me, you hadn't realised what or how to change size of SWAP partition then hopefully this will be of use!\r\n\r\nSadly for me there's probably not sufficient time left now for me to run the models..! ",
    "129099": "Thanks Jadiel, would you recomend using SGD or something like adam? Also are people using decay/momentum or just a static learning rate here? I've just reduced it to 0.0001 with SGD and it seems to actually be converging now in the first epoch.\r\n\r\nAlso, has anyone trained on batches this small and if not am I missing something, my GPU has 4GB of RAM and can't fit anything larger than a batch of 8.",
    "129095": "You need to reduce your learning rate.  With a smaller batch size the learning rate needs to also be smaller.",
    "129093": "Hi, i'm fine-tuning the vgg-16 model on an AWS g2 instance with keras & tensorflow backend. I'm only able to use a batch size of 8 before running out of memory, each epoch is also taking around 8000s which seems a lot longer than anyone else here. After 1 epoch so far I get no convergence at all, my loss has been around 14.7 since the beginning.\r\n\r\nI have pre-processed the images correctly (converted from RGB to BGR and done the VGG mean subtraction) and also converted the Theano type model to a Tensorflow type.\r\n\r\nDoes anyone have any advice, how have people managed to get this working and converging to as low as 0.3?",
    "129028": "Thanks for sharing the code and your insight!",
    "128803": "It seems that on the g2 instance, theano is only using my GPU ram (which is limited to 4gbs) but not switching to my system ram when the memory runs out. Anyone know how to fix this? Thanks!",
    "128787": "Anyone that decided to use this script with smaller images sizes (64 x 64 or 128 x 128) manage to get decent scores (< LB 0.8)? I have limited memory ~15 gb and I can only load the images at a smaller size to fit, but I'm not getting good scores (~LB 2). I messed around with ZFTurbo's original keras script and the best I could do was 0.8, so I was trying to go with the pre-trained model approach. Any help would be appreciated! Thanks!",
    "128665": "@Sandeep42\r\n I had similar problem while loading test data, and I solved it by loading and then predicting in several batches. Haven't figured out the workaround on training data yet. ",
    "128631": "A silly question maybe, but I'm suspicious, does it make sense to do the following :\r\n\r\n    train_data = np.array(train_data, dtype=np.uint8)\r\n    ...\r\n    train_data = train_data.astype('float32')\r\n\r\n, instead of doing it at once like:\r\n\r\n    train_data = np.array(train_data, dtype=np.float32)",
    "128616": "[quote=Sandeep42;128547]\r\n\r\n[quote=HeshamEraqi;126450]\r\n\r\n@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.\r\n\r\n[/quote]\r\n\r\nI have been working on this problem from a couple of days. I have the same problem as you mentioned. It seems to me that when subtracting the mean pixel from the resized image, it creates a float. My RAM is a bottleneck when images are converted into floating points. Is there any way to get around this problem?\r\n\r\n[/quote]\r\n\r\nI believe that truncating the float's into uint8's won't highly deteriorate the performance.\r\n\r\n@Ferris: The script is OK, the float-int issue arose from some custom script I've added then I fixed it. \r\n\r\n",
    "128600": "I have two questions about pre-processing of the images for a VGG model. I appreciate it if you share your thoughts. For these questions, I use the python cv2 package to read images. These questions basically ask how the original VGG model was trained.\r\n\r\n1. I see two different approaches for changing the axis of color channels in the shared scripts. Approach (a) uses the transpose method: \r\n    img = img.transpose((2,0,1))\r\nand approach (b) uses the swapaxes:\r\n   img = img.swapaxes(2, 0)\r\n   \r\nApproach (b) actually rotates the picture by 90 degrees as seen below:\r\n\r\n>>> img.shape\r\n\r\n>>>(480, 640, 3)\r\n\r\n>>> img.transpose((2, 0, 1)).shape\r\n\r\n>>> (3, 480, 640)\r\n\r\n>>> img.swapaxes(2, 0).shape\r\n\r\n>>>(3, 640, 480)\r\n\r\nWhich approach is actually correct in the sense that VGG was trained using that? My feeling is that approach a in which the image is not rotated is correct but ironically I get worse results when I use that.\r\n\r\n2. As mentioned by @Ehsan, the cv2.imread method by default reads images in BGR format. So should we also change the color space (i.e.  img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)) if VGG is trained based on RGB.  For this part, the VGG paper explicitly states RGB but this color space transformation seems to be missing from @Jiao 's code. So I am not sure if it is necessary or not. Again, when I add the transformation from BRG to RGB to my code, the performance deteriorates. \r\n\r\nI also see the same deterioration in performance when I subtract the means from each color channel. Of course, the performance deterioration may be caused by overfitting but I am still wondering what is the proper approach  based on how the original model was trained.\r\n\r\nThanks.\r\n ",
    "128547": "[quote=HeshamEraqi;126450]\r\n\r\n@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.\r\n\r\n[/quote]\r\n\r\nI have been working on this problem from a couple of days. I have the same problem as you mentioned. It seems to me that when subtracting the mean pixel from the resized image, it creates a float. My RAM is a bottleneck when images are converted into floating points. Is there any way to get around this problem?\r\n\r\n",
    "128470": "[quote=HeshamEraqi;126450]\r\n\r\n@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.\r\n\r\n[/quote]\r\n\r\nHi HeshamEraqi,\r\n\r\nI think I'm having the same problem. Can you share a tip please? Thanks.",
    "128461": "@RafayZiaMir\r\ndarknet/src/yolo.c function test_yolo",
    "128437": "@ChrisJun very clear answer, thanks a lot.",
    "128222": "Does anyone know which file(C file) we need to change in darknet to crop image? actually i cant find that file which i need to change",
    "127357": "@nanoix9\r\n\r\nmean_pixel comes from ImageNet pre-trained VGG16 model.\r\n\r\nhttps://gist.github.com/baraldilorenzo/07d7802847aaad0a35d3\r\n\r\nIn this link, you will see images are preprocessed in the order of \r\n\r\n1) Resize : Resizing image to 224x224 to fit the input shape of VGG16 model\r\n\r\n2) Subtract mean_pixel : Subtracting ImageNet mean pixels of RGB values to get zero-centered data(zero-centered data is almost always preferred in neural network)\r\n\r\nIt is a default pre-processing values when using ImageNet pre-trained model.\r\n\r\nIt comes from averaging the RGB(Red, Green, Blue) values of ImageNet data.\r\n\r\n\r\nIf you want to train data from scratch, you should use mean_pixel of you training data.\r\n\r\nChris\r\n\r\n[quote=nanoix9;127053]\r\n\r\n@Jiao Dong I saw this in the preprocessing part of code\r\n\r\n    mean_pixel = [103.939, 116.779, 123.68]\r\n    for c in range(3):\r\n        train_data[:, c, :, :] = train_data[:, c, :, :] - mean_pixel[c]\r\n\r\nwhere is the `mean_pixel` comes from? Can I use some formula instead of hard coded numbers?\r\n\r\n[/quote]\r\n",
    "127053": "@Jiao Dong I saw this in the preprocessing part of code\r\n\r\n    mean_pixel = [103.939, 116.779, 123.68]\r\n    for c in range(3):\r\n        train_data[:, c, :, :] = train_data[:, c, :, :] - mean_pixel[c]\r\n\r\nwhere is the `mean_pixel` comes from? Can I use some formula instead of hard coded numbers?",
    "126452": "Thanks for sharing. But i gave it up for my GPU is so weak. ",
    "126450": "@Ehsan & @ChrisJung: Thank you so much. I solved it. The problem was in wrong data conversion between int and floats.",
    "126374": "@ HeshamEraqi\r\nAlso, try shuffle with different seed. It should converge in first epoch.",
    "126369": "@ HeshamEraqi\r\n\r\nTry visualizing your image right before you feed into the keras model.\r\nAlso, try running your keras network in toy example (you can randomly download ~30 images from google).\r\n\r\nThe point here is to identify whether the bug lies in the input image or model construction.\r\nYou learn the most through debugging after all :)\r\n\r\nChris\r\n\r\n[quote=HeshamEraqi;126334]\r\n\r\n@xyz & @Ehsan I tried decreasing learning rate, converted RGB to BGR, and scaled images to [0-1] nothing succeeds for me and loss is stuck ~2.3. Any hints what else could be the reason ?\r\n\r\n[quote=Ehsan;126113]\r\n\r\nSkimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "126334": "@xyz & @Ehsan I tried decreasing learning rate, converted RGB to BGR, and scaled images to [0-1] nothing succeeds for me and loss is stuck ~2.3. Any hints what else could be the reason ?\r\n\r\n[quote=Ehsan;126113]\r\n\r\nSkimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "126115": "@Ehsan\r\nHi, Ehsan. Thank you for your reply. Could you mail me your email address? I have some questions and need for your help.\r\nThanks!",
    "126113": "Skimage load images in RGB format, but VGG trained on BGR format, so you need to convert RGB to BGR. To converge faster, you also need to rescale images to [0-1], even though that VGG trained on [0-255].\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n",
    "126097": "@HeshamEraqi, you may need to decrease the learning rate and try again. \r\n\r\n\r\n[quote=HeshamEraqi;126093]\r\n\r\nMy loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png\r\n\r\n[/quote]\r\n",
    "126093": "My loss is stuck around 2.3, I don't change the learning parameters in main.py. Any hints why ?\r\nI just use skimage.io and skimage.transform for imread and resize respectively instead of cv2. I also load the pretrained weights downloaded from [here][1]. I use latest Keras version.\r\n![enter image description here][2]\r\n\r\n\r\n  [1]: https://drive.google.com/file/d/0Bz7KyqmuGsilT0J5dmRCM0ROVHc/view\r\n  [2]: https://s32.postimg.org/asm8jqqmt/eeee.png",
    "125843": "Here is my update on best single model\r\nSingle googlenet-CAM can give LB 0.38746 and single VGG-CAM give LB 0.27369. \r\n\r\nsee below:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/125842#post125842",
    "125690": "I am sorry, it is a typo error. It should be 0.32. (and not 0.23)\r\n\r\n[quote=DavidGbodiOdaibo;125688]\r\n\r\n@Heng CherKeng, your vgg16 score is suspicious for a single network. If you are using the solution on this thread, your vgg16 score is based on an ensemble of 8 VGG16 models. \r\n\r\n[/quote]\r\n",
    "125688": "@Heng CherKeng, your vgg16 score is suspicious for a single network. If you are using the solution on this thread, your vgg16 score is based on an ensemble of 8 VGG16 models. ",
    "125680": "Here is my best results for single network. I use finetunning from imageNet pretrained network.\r\n\r\n - googlenet: 0.55\r\n - vgg16: 0.32 \r\n - resnet-50: cannot get it to work\r\n\r\n[quote=Ehsan;125674]\r\n\r\nDoes anybody has a good result with VGG-19 or google-net or other networks?\r\n\r\n[/quote]\r\n",
    "125679": "I have tried fine-tuning VGGnet & resnet in Torch. \r\nGot more or less the same results as you got. \r\nI expected resnet to bring better results but it was not the case.\r\n\r\n[quote=SecondPlan;125678]\r\n\r\nvgg-19: 0.20\r\n\r\nresnet:0.43\r\n\r\ngooglenet:0.53\r\n\r\n[/quote]\r\n",
    "125678": "vgg-19: 0.20\r\n\r\nresnet:0.43\r\n\r\ngooglenet:0.53",
    "125674": "Does anybody has a good result with VGG-19 or google-net or other networks?",
    "125488": "[quote=BenediktSchifferer;125427]\r\n\r\nI have the same problem like Jadiel.\r\nMy accuracy keeps at 0.10. But I use tensorflow with the pretrained model Inception_v3 from google.\r\nI may dont understand the way using pretrained models:\r\n\r\n* I load the NN structure + load the weights of the pretrained model\r\n* I delete the last layer (from last features -> classes)\r\n* I add my own classifier (from last features -> my new classes)\r\n* I process the image with the pretrained model until the last layer (so I get the feature vector)\r\n* I train my classifier from features -> my new classes\r\n\r\nIs this correct? Or do you load the weights and retrain the whole model over all layers?\r\n\r\n\r\n[/quote]\r\n\r\nafter the third step, you need to train the model and then do your 4th step ( the features you obtain is finetuned towards your dataset)\r\n",
    "125427": "I have the same problem like Jadiel.\r\nMy accuracy keeps at 0.10. But I use tensorflow with the pretrained model Inception_v3 from google.\r\nI may dont understand the way using pretrained models:\r\n\r\n* I load the NN structure + load the weights of the pretrained model\r\n* I delete the last layer (from last features -> classes)\r\n* I add my own classifier (from last features -> my new classes)\r\n* I process the image with the pretrained model until the last layer (so I get the feature vector)\r\n* I train my classifier from features -> my new classes\r\n\r\nIs this correct? Or do you load the weights and retrain the whole model over all layers?\r\n",
    "125362": "It computes bottlenecks and uses generators to be gentle with memory. Unfortunately, my script does not work well.  The loss do not diminishes with time.  It's accuracy keeps at 0.10 after 7 epochs.  What could be wrong with it? Is it because I don't use batches? Is it that I am training wrong? Anyone has an idea?",
    "125250": "@Polaris, my issue was resolved after I changed to a smaller LR. The training loss started to decrease like I expected. Did you see what the loss looked like after the 1st epoch? Did you shuffle your training data between the epochs? And you might want to  double check the way you popped the layers after loading the pretrained weights. \r\n\r\nI learned alot by doing this project, but I've moved on to something else. Based on my experience with the pretrained VGG16 model and what I see on the forum,  you are likely to get better results if you start with image size of 224x224 (default of the pretrained vgg16) like Jiao Dong mentioned in the first post, and 'learn' slowly (small LR). I couldn't do 224x224 because of memory issue, and I didn't want to keep paying more money on the Amazon EC2 instance just to try out a bigger image size.  \r\n\r\nGood luck to you.",
    "125243": "@Heng CherKeng\r\nOnce more,I sincerely thank you for your reply,that's meaningful and instructive.\r\nI will read over your words and references.\r\nIt is a \"trial and error\",thank you!",
    "125241": "@Vladimir Iglovikov\r\nI tried,but failed.\r\nCould you share the modifications you have done,such as image size and so on. \r\nThanks.\r\n[quote=Vladimir Iglovikov;125238]\r\n\r\nAfter some tweaks and modifications I got 0.19 at public LB.\r\n\r\nDid anyone try to use bigger image size and not 224x224 ?\r\n\r\n[/quote]\r\n",
    "125240": "@JennyYu\r\nYeah, I tried LR e-2, -3,even -6, however it seems useless that the training loss still stunned at 14 after one epoch.\r\nI know something I understand in a wrong way,but I just have no idea about it.\r\nDid you solved your issues after you tried different learning rate?",
    "125238": "After some tweaks and modifications I got 0.19 at public LB.\r\n\r\nDid anyone try to use bigger image size and not 224x224 ?",
    "125207": "Thanks Jiao Dong for the script. I've been able to get a 0.32 LB score with eight folds and 5 epochs. \r\n\r\nI tried reproducing the VGG16-Keras results with Caffe (using the same configuration), but was unsuccessful. I tried to pretrain   GoogleNet and AlexNet as well  ... but did not get any decent results on Caffe.  Has anyone had any success with this problem using Caffe? \r\n",
    "125202": "@ Polaris, I had similar issue as you. I used image size of 64x64 due to memory issues, and I randomly initialized the fully connected layers. My training loss was stunned due to learning rate, not because of my model. Did you try to change your LR and see what happens? ",
    "125174": "@Heng CherKeng\r\nSincere thanks for your sharing.I just wanted to re-produce your idea.But I have some troubles.\r\nWhat's the meaning of \"just change input to 256x256. the convolution filters are now applied to larger area and that's all. you still can use the pretrained vgg16 filters.\"\r\nand \"As for the last fully connect layers, yo can randomly initialised. \"\r\n\r\nSince  randomly initialized the weights,my training loss got stunned,it couldn't decrease from the beginning\r\n\r\n",
    "125093": "Thanks, which package do you use? Is that caffe?\r\n\r\nThanks,\r\n\r\n[quote=Ellen Gao Jian;125004]\r\n\r\nBelow is my result using vgg16 and batch size 16. I don't think a batch size of 16 would stop convergence.\r\n\r\n[quote=Ehsan;124802]\r\n\r\nThanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?\r\n\r\n[/quote]\r\n\r\n\r\n  [1]: http://file:///C:/Users/jgao/Desktop/1413561361361513413.JPG\r\n\r\n[/quote]\r\n",
    "125015": "@Jiao Dong:\r\n\r\nLike many other before - thank you for sharing your code. Your findings are very helpful.\r\nLike many others :) I have a question, as well... \r\n\r\nHow do you use pre-trained models? You are loading the weights from '../input/vgg16_weights.h5' and add an additional softmax layer.\r\n\r\nYour last layer has:\r\n* input-1000 features (last layer of VGG16, the 1000 classes)\r\n* output-10 features (10 classes our drivers)\r\n\r\nWhen you run model.fit() do you retrain all layers? So you retrain the weights of VGG16 and your additional layer?",
    "124802": "Thanks Heng CherKeng and Jiao Dong, I tried to train VGG16 same as you, it doesn't converge same as you. The only difference is that my batch size is 16 (due limited GPU memory). Is that make big differences?",
    "124772": "I believe that you need change the learning rate to a smaller value, such as 1e-4\r\n\r\n[quote=JennyYu;124747]\r\n\r\nI resized my image to 64 x64, and popped the last layer of hte vgg16, it's not working. My training loss is about 14, accuracy of 10%, so it's doing random guesses for the 10 category. I wonder if the vgg16 model only works well if I main the image size 224 like Jiao Dong just mentioned. Has anyone else got it to work with smaller sized image? \r\n\r\n[/quote]\r\n",
    "124747": "I resized my image to 64 x64, and popped the last layer of hte vgg16, it's not working. My training loss is about 14, accuracy of 10%, so it's doing random guesses for the 10 category. I wonder if the vgg16 model only works well if I main the image size 224 like Jiao Dong just mentioned. Has anyone else got it to work with smaller sized image? ",
    "124677": "Thanks Jiao Dong. I want to reproduce this result in Tensorflow r8.0 with GPU titanX. But the model always suffer from overfitting and the validation loss can only reach 0.8 without data augmentation and 0.44 with a lot different data augmentation. That's a surprise that you can just use the origin data to get such an amazing result.\r\n\r\nI am thinking that is it because of the precision of K40 and TitanX  or the dl tool tensorflow or theano?\r\n",
    "124670": "just change input to 256x256. the convolution filters are now applied to larger area and that's all.\r\nyou still can use the pretrained vgg16 filters.\r\n\r\nAs for the last fully connect layers, yo can randomly initialised. I tried 256x256, and now 300x300, 384x384. Will post the results when they are out.\r\n\r\n[quote=Ehsan;124631]\r\n\r\nHow do you train 256x256 with vgg16 pretrained model. The weights was computed for 224x224. Do you only use model without weights?\r\nThanks\r\n\r\n[quote=Heng CherKeng;124465]\r\n\r\ni am training a vgg with 16 layers. All layers are fine tuned.\r\nThe split of the train and test set are based on drivers.\r\nrandom 4 drivers are used as validation and remaining 22 are used for training.\r\n\r\nFrom various submissions i made, i find that:\r\n\r\n 1. So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.\r\n\r\n 2. For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.\r\n \r\n\r\n[quote=Ehsan;124440]\r\n\r\nVery interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "124669": "You are very welcome ! Glad people succeed to run it :)\r\n\r\nThis script was my first experience with keras too, thanks for ZFTurbo's original script as well that saved a lot of time. I always believed the best part of building a project is not \"take something online and everything just worked out\", it's the days & weeks you spent to debug / customize / hack it. That's how I learn from my projects too :)\r\n\r\n[quote=ChrisJung;124594]\r\n\r\nDear Jiao Dong,\r\n\r\nThank you very much for uploading the script.\r\nI also faced few Memory Errors due to my limited Hardware spec and few syntax errors due to the difference in the library version.\r\nI resolved those issue by using small batch_size, dropping shuffle by permutation and other minor adjustments for memory efficiency.\r\nSyntax error was resolved by the post in this forum!\r\n\r\nBut it was a great debugging experience, which made me learn tons! :)\r\nIt was my first touch with Keras, and fortunately I succeeded in replicating basic pipeline of your script!\r\nFew weeks of personal struggle but it was all worth it :)\r\n\r\nThanks again,\r\n\r\n[/quote]\r\n",
    "124646": "@Ehsan\r\nNo, you resize the images to 224 x 224 or use cropping. Most frameworks have already functions for that.",
    "124631": "How do you train 256x256 with vgg16 pretrained model. The weights was computed for 224x224. Do you only use model without weights?\r\nThanks\r\n\r\n[quote=Heng CherKeng;124465]\r\n\r\ni am training a vgg with 16 layers. All layers are fine tuned.\r\nThe split of the train and test set are based on drivers.\r\nrandom 4 drivers are used as validation and remaining 22 are used for training.\r\n\r\nFrom various submissions i made, i find that:\r\n\r\n 1. So long as your loss on your validation set is below 0.25, the leader board score will be quite close. Hence monitoring your validation loss is a good indicator of the leader board loss.\r\n\r\n 2. For those who find large differences between the validation and leader board scores, the accuracy of your model is probably not high enough.\r\n \r\n\r\n[quote=Ehsan;124440]\r\n\r\nVery interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "124594": "Dear Jiao Dong,\r\n\r\nThank you very much for uploading the script.\r\nI also faced few Memory Errors due to my limited Hardware spec and few syntax errors due to the difference in the library version.\r\nI resolved those issue by using small batch_size, dropping shuffle by permutation and other minor adjustments for memory efficiency.\r\nSyntax error was resolved by the post in this forum!\r\n\r\nBut it was a great debugging experience, which made me learn tons! :)\r\nIt was my first touch with Keras, and fortunately I succeeded in replicating basic pipeline of your script!\r\nFew weeks of personal struggle but it was all worth it :)\r\n\r\nThanks again,",
    "124581": "Hi, this is the first time I work with ConvNet and Keras, and I have a question about normalization. In the first post, Jiao Dong said:\r\n\r\n\"Pre-trained models are trained on ImageNet, so the normalization for pictures is a bit different; you only need to subtract the mean pixel value for each of RGB channel of a picture, instead of dividing every pixel value by 255.\"\r\n\r\nI've been researching this quite a bit, but it's still not clear to me when it's applicable to rescale by dividing by 255. For my starter code, I divided each pixel by 255, then subtracted the mean in every channel. I'll go back and just subtract the mean value without the division by 255.  But can anyone explain the different normalization methods (esp. rescaling) ? Thank you!\r\n\r\nJenny",
    "124450": "Did you have a look at the training images?\r\nMany c9 \"talking to passenger\" images look like c0 \"normal driving\" images.\r\nSo I think it's really hard or near to impossible to get this right for a NN if even a human would classify some images differently.\r\nThat's also reflected in your confusion matrix.",
    "124440": "Very interesting!\r\nDid you keep first 11 layers and train remaining layers?\r\nDid you split validation based on drivers or just random?\r\n\r\nThank you",
    "124121": "Thanks ~ so glad to see it worked on other server lol\r\n\r\nYes i used cuda 7.5 and have cudnn installed in server environment. Also since I am using Theano, in configuration file I also enabled fastmath and cnmem.\r\n\r\n[quote=scsherm;124118]\r\n\r\nThe code works and all the information is in this forum. If you are using a newer version of keras, make sure to use remove the last layer with the code below. Also, I am using a gpu instance on aws. Do you have  a fast gpu on you local machine? Did you install cuDNN and cuda?\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\n[/quote]\r\n",
    "124092": "So finally is there anyone manage to get the code worked? I have tried but still can't run on my local machine.",
    "123700": "You are right, it's the public leaderboard score.\r\n\r\nConsidering log score is computed by a sum over log confidence of all pictures, in our own cross validation we would only have about ~3000 pictures, but in testing dataset there are ~79,000. So the LB score is definitely much higher than training loss score, simply because of the size of dataset.\r\n\r\n[quote=Duc Nguyen;123697]\r\n\r\nAre those the loss scores in your training, or scores in the public leaderboard? \r\nI guess the later since you had much lower validation loss. \r\nIn this case, I think I am missing something in my Torch code.\r\nThanks again.\r\n\r\n[quote=Jiao Dong;123690]\r\n\r\nAt 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "123697": "Are those the loss scores in your training, or scores in the public leaderboard? \r\nI guess the later since you had much lower validation loss. \r\nIn this case, I think I am missing something in my Torch code.\r\nThanks again.\r\n\r\n[quote=Jiao Dong;123690]\r\n\r\nAt 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "123691": "good call :)\r\n\r\n[quote=Vinh Nguyen;122366]\r\n\r\nfor the newer version of Keras, you need a slightly different code to do model surgery, i.e. pop out the last layer and insert the new layer\r\n\r\n    model.layers.pop()\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n    model.add(Dense(10, activation='softmax'))\r\n\r\n\r\n\r\n[quote=Jiao Dong;121903]\r\n\r\n@Manuele Tamburrano\r\n\r\nI think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)\r\n\r\nBecause as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.\r\n\r\nI did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(\r\n\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "123690": "At 15 epochs each model itself has nearly identical loss score within a small range 0.24~0.28 I would say, and model ensemble tend to yield better result than individual model.\r\n\r\n[quote=Duc Nguyen;123685]\r\n\r\n@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n\r\n\r\n[/quote]\r\n",
    "123685": "@Jiao Dong Thanks for sharing the code.\r\nDo you have any test result using just a single model? I am curious about the performance of your single best model, trained without data augmentation.\r\nI am trying to train a CNN using Torch. However, if I skip the data augmentation and simply use 224x224 images, the network overfit very quickly, even with high dropout ratio. \r\n",
    "123554": "Can I confirm if my understanding for the  LB 0.23800 solution is correct or not?\r\n\r\n -  No data augmentation was used. The image was just simply resized to 224x224.\r\n      Color image was used.\r\n \r\n - The results is the average of 8 models. The train data was divided into 8 folds. For training each model,\r\n      one fold served as validation set and the remaining 7 as training set. The validation set was used \r\n      to determine when to terminate the training.\r\n\r\nThanks!",
    "122349": "@Jiao Dong Sorry for my previous post, but the fact remains that some people were a little confused. Your code was very helpful for many people:-)",
    "121963": "From my understanding of the data, using validation_split = 0.2 is the right way of solving the problem, which is likely to produce the optimize PB score.\r\n\r\nThe splitting base on the drivers just give an estimation of how generation the model is, i.e your score should be expected to have the average of the K-fold valid loss. It does not guarantee producing the best models in this case. \r\n",
    "121939": "@Jiao Dong\r\nOh, I didn't catch that haha. Too bad that my computer is having a memory error with loading the data in RGB, so I can't even tryout your code... :(  Anyway, thx again for the code! Hope you aced your project :) ",
    "121904": "@Jeong Wook Moon\r\n\r\nIn the function signature if a parameter is not explicitly passed it would use the default value, like:\r\n\r\ndef load_train(img_rows, img_cols, color_type=1):\r\n\r\nwould use grayscale images if i didn't specify to use color images.\r\n\r\nAt the part right after imports, there's a global variable \r\n\r\ncolor_type_global = 3\r\n\r\nwhich is the parameter actually passed into \r\n\r\n    train_data, train_target, driver_id, unique_drivers = \\\r\n        read_and_normalize_and_shuffle_train_data(img_rows, img_cols,\r\n                                                  color_type_global)\r\n\r\nmodel = vgg_std16_model(img_rows, img_cols, color_type_global)\r\n\r\ntest_data, test_id = read_and_normalize_test_data(img_rows, img_cols,\r\n                                                      color_type_global)\r\n\r\ntest_data, test_id = read_and_normalize_test_data(img_rows, img_cols,\r\n                                                      color_type_global)\r\n\r\n\r\nDid this help for your question ? :)",
    "121903": "@Manuele Tamburrano\r\n\r\nI think the first thing I should do is to verify we are using the exact same setup, like same repo of libraries with same version. (keras, anaconda, theano, cuda, cudnn , etc.) If my files are still there hopefully I can post it later today :)\r\n\r\nBecause as you pointed out before, I forgot to mention when I accidentally used a newer version of keras from github then same code did not even seem to converge for some reason. There might be more issues like that i didn't test on other servers.\r\n\r\nI did my split base on drivers for the first couple runs, but then I started to shuffle training images before and between each epoch, that's how I got my LB 0.32 ~ 0.238 submissions. Since it was the last couple days before deadline I didn't get the chance to run the same parameters based on drivers, I can't say for sure if it would better or not :(\r\n\r\n",
    "121901": "@rcarson @Manuele Tamburrano\r\n\r\nSorry what i wrote previously is not towards you guys :) , it's the other posts i just saw which seemed to claim I posted the wrong information on purpose to mislead people, especially I didn't think they even tried to setup the environment & code to execute it themselves. \r\n\r\nI'm glad you have been trying to execute the code and debug, I guess due to the different versions of libraries and frameworks it might behave a bit differently, that's something I didn't consider when i post it because I was assuming as long as the script is the same it would yield similar result, which seems to be wrong :(\r\n\r\nThe admin I have been trying to contact is on vacation ... but I would get the server up and running as long as he comes back and see what's the issue.  \r\n\r\n",
    "121897": "[quote=Jiao Dong;121891]\r\n\r\nI have moved on to my other priorities after this post, but it seemed people still have questions about the script, and it's a bit annoying since when i was confused I usually run experiments on my own, do some research and ask specific questions before assuming its false.\r\n\r\nI will contact my school's admin to restore my directory to the state before last semester ends, and probably upload a video of its execution, I would also go through the script source code line by line before it runs and show it's the same as the one I posted. \r\n\r\nI spent couple hours to write this post in a monetary competition because I got help from other people's posts as well. But I didn't expect to waste couple more hours to show it works to convince people who didn't even try to debug on their own.\r\n\r\n\r\n\r\n[/quote]\r\n\r\nI never doubted about your code, if you read my comments I wonder if I did something wrong or if my device is different for some reason (float precision) to reproduce your results. I'm sorry about unpleasant comments of users, I can understand they hurt when you share something very useful in a paid competition. Thank you again for your code.\r\n\r\nAnyway I spent several hours before giving up to go below 0.4 on LB score, I tried several minor changes on your code, used driver split to detect overfitting, changed parameters, kfold methods, but still stuck on ~0.4.\r\n\r\nNow I'm on different road, but I'm still curious on what is the cause to this difference.\r\n@rcarson, do you use same frameworks and versions of Jiao (anaconda, keras, theano) or the more updated ones? Are you splitting the train by driver or just a random split?\r\nFrom my tries splitting by drivers results in a better estimate of overfitting, but the LB score is worse\r\n",
    "120491": "anyone succeded to reproduce that LB score using these code without tweaking parameters?\r\n\r\nI did several tries with changes proposed by me and jvipond  and adjusting learning rate but I cannot go beyond ~0.4 on LB score.\r\n\r\nI'm moving on different solutions but I'm still wondering why the difference is so big running the same code (probably with newer versions of libraries). So I'm a bit worried that there is something shady in my tries, like float precision or something like that.\r\n\r\nSo, anyone was able to get better results with the same code?",
    "119657": "@ Jiao Dong\r\nThank for your details explanation and sharing! I would like to ask have you tried vgg-19 with Keras?",
    "119387": "I'm trying a different approach, I edited the load_weights method in engine/topology.py to only load layers originally defined. So I can simply remove the dense layer with 1000 neurons and add a new one with only 10 neurons without having to pop anything.\r\nIn this way parameters are 40970 that seems correct, the loss is decreasing faster but apparently not as fast as op's loss. I'll let the net to train this night to see where it is going",
    "119384": "I think @jvipond got the solution for the problem, the way you change dense layer is different  in newer versions of Keras. I also spent quite a bit of time messing with libraries on the remote server so my setup might be different as well.\r\n\r\nI can take a look of the layers later, but got to finish finals first =__=\r\n\r\n[quote=Manuele Tamburrano;119377]\r\n\r\nthank you for clarifications, but probably you are remembering something wrong.\r\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?\r\n\r\nAnd still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\r\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons\r\n\r\n[quote=Jiao Dong;119371]\r\n\r\nI don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here https://anaconda.org/omnia/keras (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.\r\n\r\nYou reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. \r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "119377": "thank you for clarifications, but probably you are remembering something wrong.\r\nKeras 0.8.0 should not exist as far as I know, maybe are you referring to Theano version?\r\n\r\nAnd still as someone pointed, the weights seems wrong, are you able to load your model and print the last two lines of model.summary() method output?\r\nParams are 10010, so probably when you pop the layer, params are not popped and you end with a dense layer with 1000 neurons attached to a layer with 10 neurons\r\n\r\n[quote=Jiao Dong;119371]\r\n\r\nI don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here https://anaconda.org/omnia/keras (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.\r\n\r\nYou reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. \r\n\r\n\r\n[/quote]\r\n",
    "119374": "Yep, just need to change prototxt file for the output layer, and add a data layer on top to feed into the model.\r\n\r\nWe ran out of time to do more experiments before deadline (May 1st) so we didn't try anything further. But seriously, we are using the same GPU but training on his Caffe is like 3~5 times faster than mine\r\n\r\n[quote=Bojan Tunguz;119372]\r\n\r\n[quote=Jiao Dong;119367]\r\n\r\nMy teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.\r\n\r\n\r\n[/quote]\r\n\r\nIs that with the same data augmentation and CV as done in this Keras + VGG_16 script?\r\n\r\n\r\n[/quote]\r\n",
    "119372": "[quote=Jiao Dong;119367]\r\n\r\nMy teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.\r\n\r\n\r\n[/quote]\r\n\r\nIs that with the same data augmentation and CV as done in this Keras + VGG_16 script?\r\n",
    "119371": "I don't have root access to the server I used so I installed Anaconda on my local directory, then I installed my Keras from here https://anaconda.org/omnia/keras (Somehow they changed version for this package, but as I can remember it should be 0.8.0), all other libraries are installed through conda as well.\r\n\r\nYou reminded me of a similar situation; I updated Keras (1.0.1?) from github installation to test TensorFlow, but I didn't change it back. Then for the same code (pretty much the same as this one), training loss seems would fail to converge for some reason, but later when I use the 0.8.0 version on anaconda cloud, it behaves normally. \r\n\r\n[quote=Manuele Tamburrano;119370]\r\n\r\nyes, I tried to lower learning rate and some small adjustment but the loss keeps decreasing very slowly, I get ~1.3 after 15 epochs.\r\n\r\nWhat version of Keras are you running? The one from pip (it should be 1.0.2) or master compiled from source? If this is the case, what's your last commit hash?\r\n\r\nSeems like you are using an old version, because your logs show accuracy values, but you only set \"show_accuracy=True\" in the fit method, but in the last version you should use metrics=[\"accuracy\"] in compile method\r\n\r\n[quote=Jiao Dong;119366]\r\n\r\nThere's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.\r\n\r\nFrom replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?\r\n\r\n\r\n[/quote]\r\n\r\n\r\n[/quote]\r\n",
    "119367": "My teammate setup Caffe on this computer running ResNet-50, he got 0.25374. It's also amazingly fast and flexible compare to my script using keras.\r\n\r\n[quote=Matteo Presutto;119360]\r\n\r\nDid any of you try fine tuning resnet50? The best I can get out of it is 0.85, It starts overfitting after 2-3 iterations\r\n\r\n[/quote]\r\n",
    "119366": "There's a small chance it might be, because I have been making small changes to this file after my submission as well.. But overall this script includes all the changes for the 0.23800 loss submission and I've written all the details.\r\n\r\nFrom replies in this thread it seemed people would get different training loss based on the same script.  In my case, I remember my training loss started from 4~5 and converged to something below 1 in first 10,000 pictures, first epoch. It might behave differently on another machine based on your setup.   Did you try to change parameters for your model, like learning rate ?\r\n\r\n[quote=Manuele Tamburrano;119348]\r\n\r\nThank you for sharing this code, I was trying to do essentially the same.\r\n\r\nAnyway I can't reproduce your results, my loss decreases very slowly and I can't see the fast improvements you talk about in the first ~1000 to ~5000 images in the first epoch.\r\n\r\nDid you change anything else but the num_fold and nb_epochs compared to the published code?\r\n\r\n[/quote]\r\n",
    "119360": "Did any of you try fine tuning resnet50? The best I can get out of it is 0.85, It starts overfitting after 2-3 iterations",
    "119348": "Thank you for sharing this code, I was trying to do essentially the same.\r\n\r\nAnyway I can't reproduce your results, my loss decreases very slowly and I can't see the fast improvements you talk about in the first ~1000 to ~5000 images in the first epoch.\r\n\r\nDid you change anything else but the num_fold and nb_epochs compared to the published code?",
    "119186": "LOL it shows how much difference between GPU and CPU in floating point computation  \r\n\r\n[quote=Gerard Toonstra;119149]\r\n\r\nI posted about how to use a memory mapped file in this thread here:\r\n\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20664/data-can-t-fit-in-memory/119147#post119147\r\n\r\n\r\nI also tried to run the network on a CPU, because my GPU doesn't have enough memory. So far, I only have one thread that's working on the data. Estimated time to finish a single epoch for the first fold:\r\n\r\n95 days.\r\n\r\nLOL! :)\r\n\r\n\r\n[/quote]\r\n",
    "119158": "[quote=Wendy Kan;119155]\r\n\r\nSince it's non-commercial use only, State Farm won't be able to use it. So the usage of VGG-16 is not allowed. \r\n\r\n  [1]: https://gist.github.com/ksimonyan/211839e770f7b538e2d8#file-readme-md\r\n\r\n[/quote]\r\n\r\nWait a sec, VGG-16 has already been used in several recent Kaggle competitions. Will you retroactively re-evaluate all of those???",
    "119157": "Oh, that's a pity... :(",
    "119140": "[quote=ZFTurbo;119102]\r\n\r\n**Abhijay Arora**, If use this code \"as is\" it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.\r\n\r\nhttp://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\r\n\r\nTrue.But with a generator it is possible it seems.Do you have a generator that can be used with this?. I tried with a simple generator. it's taking much time.I'm not sure about it's success.\r\n\r\ndef generate():\r\n\r\n    for i in range(2803):  # 2803*8 = 22424  ---> this generator returns as 8 chunks\r\n        #as i'm trying with gray scale i have it converted to array in memory (My laptop RAM 8GB)\r\n       #you may need to convert the training set to array \"here itself\"\r\n        yield (np.array(X_train[i*8:(i+1)*8],dtype=np.uint8),np.array(Y_newtrain[i*8:(i+1)*8],dtype=np.uint8))\r\n\r\nfor e in range(nb_epoch):\r\n    print(\"epoch %d\" % e)\r\n    for X_batch, Y_batch in generate(): \r\n        model.fit(X_batch, Y_batch,nb_epoch=1,show_accuracy=True,shuffle=True)\r\n\r\nPlease help if someone do have a generator.Thanks for the help in advance\r\n\r\n[/quote]\r\n",
    "119107": "@ Abhijay\r\n\r\nMy immediate guess is that this is caused by the Fully Connected layer (which expects an input at a fixed size) Changing the original input image size will mess up with this layer.",
    "119106": "[quote=ZFTurbo;119102]\r\n\r\n**Abhijay Arora**, If use this code \"as is\" it requires at least 32 GB of RAM. To use only training you can fit in 16 GB. To run this code on low RAM machine you need to fully rewrite reading part. You need to read image by small parts required by batch training. And the same for test images.\r\n\r\nhttp://keras.io/getting-started/faq/#how-can-i-use-keras-with-datasets-that-dont-fit-in-memory\r\n\r\n[/quote]\r\n\r\nThanks. I tried running the script \"as is\", just changed the image size to 64 x 64 . I get the following error while loading the weights:\r\n\r\n    Exception: Layer weight shape (2048, 4096) not compatible with provided weight shape (25088, 4096)\r\n\r\nAny idea why?\r\n",
    "119104": "It is interesting that you managed to achieve such a good score without overfitting to the persons in the training set without having any holdout validation set with persons unseen during training. Was it just alot of trial and error of parameters? ",
    "119071": "@Jiao\r\n\r\nJust wanted to quickly thank you for posting the code and for such a very thorough and detailed writeup. I have a few general questions, but I'll leave those for later after I've played with your code a bit. ",
    "119062": "Hi Jiao Dong,\r\n\r\nCan you share the parameters of you model like vgg16? Without gpu, it will take several days for 20 epoch. Or just share the visualization of first layer. Very curious about how the first layer will look like. \r\n\r\nBTW, is this a bug in? I see \"train_drivers\" and \"test_drivers\" never be used in the loop:\r\n\r\nfor train_drivers, test_drivers in kf:\r\n        num_fold += 1\r\n        print('Start KFold number {} from {}'.format(num_fold, nfolds))\r\n        # print('Split train: ', len(X_train), len(Y_train))\r\n        # print('Split valid: ', len(X_valid), len(Y_valid))\r\n        # print('Train drivers: ', unique_list_train)\r\n        # print('Test drivers: ', unique_list_valid)\r\n        # model = create_model_v1(img_rows, img_cols, color_type_global)\r\n        # model = vgg_bn_model(img_rows, img_cols, color_type_global)\r\n        model = vgg_std16_model(img_rows, img_cols, color_type_global)\r\n\r\n        model.fit(train_data, train_target, batch_size=batch_size,\r\n                  nb_epoch=nb_epoch,\r\n                  show_accuracy=True, verbose=1,\r\n                  validation_split=split, shuffle=True)\r\n\r\n\r\n",
    "119009": "That's interesting to see, unfortunately I don't know the solution either. This script is one of the many I've tried that converged well, if you discovered any insight of this model / script, please keep me updated :) \r\n\r\n[quote=threecourse;118999]\r\n\r\n@rcarson\r\n\r\nIn my case, train loss decreased slowly, around 1.8 at 5 epoch. (I quitted there) \r\nVal_loss is not reliable due to split by image not driver.\r\n\r\nIn addition, by printing model.summary(), number of weights of last layer is 10010.\r\nIt indicates layers are not connected properly. \r\n\r\n\r\n[/quote]\r\n",
    "119007": "Yep, I didn't use driver information in training and shuffled my training data on purpose. \r\n\r\nHowever I didn't run any model based on driver information with high epoch.\r\n\r\n[quote=ZFTurbo;118979]\r\n\r\nI think on these lines you actually lost the connection between drivers and images:\r\n\r\n    perm = permutation(len(train_target))\r\n    train_data = train_data[perm]\r\n    train_target = train_target[perm]\r\n\r\n[/quote]\r\n",
    "119000": "Thank you. I'll let you know if I run into the same thing or not",
    "118994": "@threecourse, hi, i'm running it now, what do you mean by 'doesn't work'? CV score is too bad?\r\n\r\nBest,",
    "118985": "It didn't work for me, anyone success?\r\n\r\nI think in the script CV is split by image, not by driver.",
    "118952": "@saihttam\r\n\r\nTry reducing the batch size",
    "118920": "[quote=Jiao Dong;118880]...it took more than twice as much time to train an epoch in TensorFlow - 2 GPU compare to Theano - 1 GPU[/quote]\r\n\r\nIs this with TensorFlow 0.8? There are supposed to be speed improvements in the latest version.\r\n\r\nMany thanks for posting your script, looking forward to trying it out!\r\n",
    "118917": "@Jiao, thanks for your sharing. This is really awesome!",
    "118946": "",
    "2008185": "thanks for explaning  this simple VGG16",
    "125776": "Thanks for sharing！",
    "124829": "Thanks for the suggestion, Yu Hai",
    "124675": "Thank you, Jiao Dong"
  }
}