{
  "id": 22017,
  "title": "End-to-end workflow using MXnet (inception BN + VGG ensemble)",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22017",
  "author_name": "",
  "post_date": "2016-07-02T22:38:52.103Z",
  "votes": 15,
  "comment_count": 7,
  "views": 3602,
  "content": "<p>Acknowledgement:  Thanks to Tianqi, Bing, Terry and Hongliang for the code work. MXnet repo <a href=\"https://github.com/dmlc/mxnet\">https://github.com/dmlc/mxnet</a> please star and fork this great repo. Special thanks to ysfalo for his fine tune blog <a href=\"https://ysfalo.github.io/2016/04/01/mxnet之fine-tune/\">https://ysfalo.github.io/2016/04/01/mxnet%E4%B9%8Bfine-tune/</a> </p>\n\n<p><strong>1. Github link <a href=\"https://github.com/phunterlau/kaggle_statefarm\">https://github.com/phunterlau/kaggle_statefarm</a></strong></p>\n\n<p><strong>2. How to use</strong></p>\n\n<p>a. Install mxnet by following the official instruction for compiling and installing python packages. Please install necessary CUDA packages before compiling. Please turn on the CUDA option, cuDNN option for fast training (cuDNN v5 is suggested), and the most important, OPENCV option for having &#8216;im2rec&#8217; binary.</p>\n\n<p>b. Run Inception-BN model</p>\n\n<p>1) Go to inception directory, download the pretrained Inception BN model from dmlc.ml by executing <code>download.sh</code>. The pretrained model will appear in <code>model</code> directory.</p>\n\n<p>2) Generate <code>.rec</code> format train-val dataset using <code>im2rec</code>: put the training dataset into <code>train</code> directory, edit <code>gen_rec_bin.sh</code> and replace <code>im2rec_path</code> variable to the location of  <code>im2rec</code> binary. Usually it is in the <code>bin</code> directory after compiling MXNet. Run gen_rec_bin.sh and generate 5 cross-validation datasets for 5 batches of training.</p>\n\n<p>3) Train the model: run <code>run.cv_inception_bn.sh</code> to execute the model training on 5 cross-validation datasets. The code will save each epoch snapshot in <code>model</code> directory, and the <code>mean.bin</code> for each batch. Please save these snapshots and <code>mean.bin</code> for testing. Usually in the first couple of epochs, the training and validation accuracy can increase and reach almost 1.000 after some epochs. By default, this script asks for 30 epochs. Each epoch takes about 150 (Titan X) - 300+ (GTX 960) seconds depending on the video card, so usually 30 epochs for all 5 folds can take around 7 hours on a single decent card with less than 2GB memory. If you have multiple video cards, please use <code>--gpus</code> options to use multiple cards, and increase batch size if you want.</p>\n\n<p>4) Inference: This is time consuming part because there are about 80k test images. In the inception directory, <code>run.test.sh</code> can load the model snapshots for each CV batch, run a center crop of test image, and produce the prediction in the directory, like <code>submission_test_batch_5_epoch_20.csv</code>. These output can be merged with other models, or be averaged for a single model submission. The single model inception-BN with the current setup can give 0.59 in the public LB.</p>\n\n<p>c. Run VGG model.</p>\n\n<p>1) Go to VGG directory, download the pretrained VGG models by executing <code>download.sh</code></p>\n\n<p>2) Generate <code>.rec</code> format train-val dataset like you have done for inception BN model. Please note that, VGG and Inception BN use different input size.</p>\n\n<p>3) Train the model. Run <code>run.cv.sh</code> for run a 5 fold 15 epochs VGG model. You might notice <code>--finetune-lr-scale 5</code> parameter in the shell script, that is because we have experimented 5x learning rate can quickly produce some good results. Training time for each epoch is about 450 (Titan X) -1000 (GTX 960) seconds, and the total training time for a 5-fold run 15 epoch each can be around 10 hours.</p>\n\n<p>4) Inference: it is similar as in Inception BN model. With the current setup, the single VGG model can give about 0.56 scores in the public LB.</p>\n\n<p>Ensemble different model output: please go to <code>merge</code> directory, make two directories <code>vgg</code> and <code>inception</code>, put the two set of output to the corresponding directories, run <code>merge.py</code> and the output is ready to be submitted. The equal weight ensemble of an inception BN model with VGG can reproduce approximately 0.51 LB.</p>\n\n<p><strong>3. How pretrain model works:</strong></p>\n\n<p>In MXnet, we can load a model and keep training with new data. In this problem, we load models from large dataset and continue to train with state farm dataset, that requires that the new data has little impact to the first two layers of the convolutional network, so the network can keep the original weights for the first two CNN layers, and learn from the new dataset for the high level features. In inception model and vgg model&#8217;s symbol python file, one can find these convolutional layers having <code>attr={'lr_mult': '0.00'}</code> in the layer description. In VGG model, using average pooling instead of two fully connected layer before classifier, because the input size as 320*320 instead of 224*224  for the original VGG model.  </p>\n\n<p><strong>4. What parameters can be experimented for improving the score</strong></p>\n\n<p>a.  Data augment: train.py has multiple augment options, and one can experiment with mx.io.ImageRecordIter, for example, random rotation, random illumination, and other options. These options can improve the LB scores by generalizing the model. Please refer to MXnet document for how to use them.</p>\n\n<p>b.  CV: this competition has very small dataset which gives large variance for models. Having multiple runs on multiple different split on train-val dataset can significantly reduce the variance. We provide 5-fold CV, and we have experiment 10-fold CV and other CV tries. Ensembling multiple CV choices can also help smoothing down the variance. Some one has asked if spliting by driver info can help, our experiment shows it is an dangerous choice and it may overfit to some unnecessary features.</p>\n\n<p>c.  Learning rate change: usually pretrained model needs a large learning rate from new data for quickly converging, so in the train_model.py we have a fine tune learning rate scale parameter <code>--finetune-lr-scale</code> for tuning up. There are other parameters like <code>--lr-factor-epoch</code> for fine tune the parameters for different epochs.</p>\n\n<p>d.  Crop in test set: in the vgg directory, one can also find <code>run.test_10crop.sh</code> for running a random crop for 10 times and  average on each test image. It can improve LB score from about 0.05 to 0.1, but it takes much longer time. One can modify and get some other crop methods, e.g. 3 crop etc.</p>",
  "messages": [
    {
      "id": "125810",
      "postDate": "07/02/2016 22:38:52",
      "content": "<p>Acknowledgement:  Thanks to Tianqi, Bing, Terry and Hongliang for the code work. MXnet repo <a href=\"https://github.com/dmlc/mxnet\">https://github.com/dmlc/mxnet</a> please star and fork this great repo. Special thanks to ysfalo for his fine tune blog <a href=\"https://ysfalo.github.io/2016/04/01/mxnet之fine-tune/\">https://ysfalo.github.io/2016/04/01/mxnet%E4%B9%8Bfine-tune/</a> </p>\n\n<p><strong>1. Github link <a href=\"https://github.com/phunterlau/kaggle_statefarm\">https://github.com/phunterlau/kaggle_statefarm</a></strong></p>\n\n<p><strong>2. How to use</strong></p>\n\n<p>a. Install mxnet by following the official instruction for compiling and installing python packages. Please install necessary CUDA packages before compiling. Please turn on the CUDA option, cuDNN option for fast training (cuDNN v5 is suggested), and the most important, OPENCV option for having &#8216;im2rec&#8217; binary.</p>\n\n<p>b. Run Inception-BN model</p>\n\n<p>1) Go to inception directory, download the pretrained Inception BN model from dmlc.ml by executing <code>download.sh</code>. The pretrained model will appear in <code>model</code> directory.</p>\n\n<p>2) Generate <code>.rec</code> format train-val dataset using <code>im2rec</code>: put the training dataset into <code>train</code> directory, edit <code>gen_rec_bin.sh</code> and replace <code>im2rec_path</code> variable to the location of  <code>im2rec</code> binary. Usually it is in the <code>bin</code> directory after compiling MXNet. Run gen_rec_bin.sh and generate 5 cross-validation datasets for 5 batches of training.</p>\n\n<p>3) Train the model: run <code>run.cv_inception_bn.sh</code> to execute the model training on 5 cross-validation datasets. The code will save each epoch snapshot in <code>model</code> directory, and the <code>mean.bin</code> for each batch. Please save these snapshots and <code>mean.bin</code> for testing. Usually in the first couple of epochs, the training and validation accuracy can increase and reach almost 1.000 after some epochs. By default, this script asks for 30 epochs. Each epoch takes about 150 (Titan X) - 300+ (GTX 960) seconds depending on the video card, so usually 30 epochs for all 5 folds can take around 7 hours on a single decent card with less than 2GB memory. If you have multiple video cards, please use <code>--gpus</code> options to use multiple cards, and increase batch size if you want.</p>\n\n<p>4) Inference: This is time consuming part because there are about 80k test images. In the inception directory, <code>run.test.sh</code> can load the model snapshots for each CV batch, run a center crop of test image, and produce the prediction in the directory, like <code>submission_test_batch_5_epoch_20.csv</code>. These output can be merged with other models, or be averaged for a single model submission. The single model inception-BN with the current setup can give 0.59 in the public LB.</p>\n\n<p>c. Run VGG model.</p>\n\n<p>1) Go to VGG directory, download the pretrained VGG models by executing <code>download.sh</code></p>\n\n<p>2) Generate <code>.rec</code> format train-val dataset like you have done for inception BN model. Please note that, VGG and Inception BN use different input size.</p>\n\n<p>3) Train the model. Run <code>run.cv.sh</code> for run a 5 fold 15 epochs VGG model. You might notice <code>--finetune-lr-scale 5</code> parameter in the shell script, that is because we have experimented 5x learning rate can quickly produce some good results. Training time for each epoch is about 450 (Titan X) -1000 (GTX 960) seconds, and the total training time for a 5-fold run 15 epoch each can be around 10 hours.</p>\n\n<p>4) Inference: it is similar as in Inception BN model. With the current setup, the single VGG model can give about 0.56 scores in the public LB.</p>\n\n<p>Ensemble different model output: please go to <code>merge</code> directory, make two directories <code>vgg</code> and <code>inception</code>, put the two set of output to the corresponding directories, run <code>merge.py</code> and the output is ready to be submitted. The equal weight ensemble of an inception BN model with VGG can reproduce approximately 0.51 LB.</p>\n\n<p><strong>3. How pretrain model works:</strong></p>\n\n<p>In MXnet, we can load a model and keep training with new data. In this problem, we load models from large dataset and continue to train with state farm dataset, that requires that the new data has little impact to the first two layers of the convolutional network, so the network can keep the original weights for the first two CNN layers, and learn from the new dataset for the high level features. In inception model and vgg model&#8217;s symbol python file, one can find these convolutional layers having <code>attr={'lr_mult': '0.00'}</code> in the layer description. In VGG model, using average pooling instead of two fully connected layer before classifier, because the input size as 320*320 instead of 224*224  for the original VGG model.  </p>\n\n<p><strong>4. What parameters can be experimented for improving the score</strong></p>\n\n<p>a.  Data augment: train.py has multiple augment options, and one can experiment with mx.io.ImageRecordIter, for example, random rotation, random illumination, and other options. These options can improve the LB scores by generalizing the model. Please refer to MXnet document for how to use them.</p>\n\n<p>b.  CV: this competition has very small dataset which gives large variance for models. Having multiple runs on multiple different split on train-val dataset can significantly reduce the variance. We provide 5-fold CV, and we have experiment 10-fold CV and other CV tries. Ensembling multiple CV choices can also help smoothing down the variance. Some one has asked if spliting by driver info can help, our experiment shows it is an dangerous choice and it may overfit to some unnecessary features.</p>\n\n<p>c.  Learning rate change: usually pretrained model needs a large learning rate from new data for quickly converging, so in the train_model.py we have a fine tune learning rate scale parameter <code>--finetune-lr-scale</code> for tuning up. There are other parameters like <code>--lr-factor-epoch</code> for fine tune the parameters for different epochs.</p>\n\n<p>d.  Crop in test set: in the vgg directory, one can also find <code>run.test_10crop.sh</code> for running a random crop for 10 times and  average on each test image. It can improve LB score from about 0.05 to 0.1, but it takes much longer time. One can modify and get some other crop methods, e.g. 3 crop etc.</p>",
      "rawMarkdown": "Acknowledgement:  Thanks to Tianqi, Bing, Terry and Hongliang for the code work. MXnet repo https://github.com/dmlc/mxnet please star and fork this great repo. Special thanks to ysfalo for his fine tune blog https://ysfalo.github.io/2016/04/01/mxnet%E4%B9%8Bfine-tune/ \r\n\r\n **1. Github link https://github.com/phunterlau/kaggle_statefarm**\r\n\r\n **2. How to use**\r\n\r\na. Install mxnet by following the official instruction for compiling and installing python packages. Please install necessary CUDA packages before compiling. Please turn on the CUDA option, cuDNN option for fast training (cuDNN v5 is suggested), and the most important, OPENCV option for having ‘im2rec’ binary.\r\n\r\n\r\nb. Run Inception-BN model\r\n\r\n1) Go to inception directory, download the pretrained Inception BN model from dmlc.ml by executing `download.sh`. The pretrained model will appear in `model` directory.\r\n\r\n2) Generate `.rec` format train-val dataset using `im2rec`: put the training dataset into `train` directory, edit `gen_rec_bin.sh` and replace `im2rec_path` variable to the location of  `im2rec` binary. Usually it is in the `bin` directory after compiling MXNet. Run gen_rec_bin.sh and generate 5 cross-validation datasets for 5 batches of training.\r\n\r\n\r\n3) Train the model: run `run.cv_inception_bn.sh` to execute the model training on 5 cross-validation datasets. The code will save each epoch snapshot in `model` directory, and the `mean.bin` for each batch. Please save these snapshots and `mean.bin` for testing. Usually in the first couple of epochs, the training and validation accuracy can increase and reach almost 1.000 after some epochs. By default, this script asks for 30 epochs. Each epoch takes about 150 (Titan X) - 300+ (GTX 960) seconds depending on the video card, so usually 30 epochs for all 5 folds can take around 7 hours on a single decent card with less than 2GB memory. If you have multiple video cards, please use `--gpus` options to use multiple cards, and increase batch size if you want.\r\n\r\n\r\n4) Inference: This is time consuming part because there are about 80k test images. In the inception directory, `run.test.sh` can load the model snapshots for each CV batch, run a center crop of test image, and produce the prediction in the directory, like `submission_test_batch_5_epoch_20.csv`. These output can be merged with other models, or be averaged for a single model submission. The single model inception-BN with the current setup can give 0.59 in the public LB.\r\n\r\n\r\nc. Run VGG model.\r\n\r\n1) Go to VGG directory, download the pretrained VGG models by executing `download.sh`\r\n\r\n2) Generate `.rec` format train-val dataset like you have done for inception BN model. Please note that, VGG and Inception BN use different input size.\r\n\r\n3) Train the model. Run `run.cv.sh` for run a 5 fold 15 epochs VGG model. You might notice `--finetune-lr-scale 5` parameter in the shell script, that is because we have experimented 5x learning rate can quickly produce some good results. Training time for each epoch is about 450 (Titan X) -1000 (GTX 960) seconds, and the total training time for a 5-fold run 15 epoch each can be around 10 hours.\r\n\r\n4) Inference: it is similar as in Inception BN model. With the current setup, the single VGG model can give about 0.56 scores in the public LB.\r\n\r\n\r\nEnsemble different model output: please go to `merge` directory, make two directories `vgg` and `inception`, put the two set of output to the corresponding directories, run `merge.py` and the output is ready to be submitted. The equal weight ensemble of an inception BN model with VGG can reproduce approximately 0.51 LB.\r\n\r\n\r\n**3. How pretrain model works:**\r\n\r\nIn MXnet, we can load a model and keep training with new data. In this problem, we load models from large dataset and continue to train with state farm dataset, that requires that the new data has little impact to the first two layers of the convolutional network, so the network can keep the original weights for the first two CNN layers, and learn from the new dataset for the high level features. In inception model and vgg model’s symbol python file, one can find these convolutional layers having `attr={'lr_mult': '0.00'}` in the layer description. In VGG model, using average pooling instead of two fully connected layer before classifier, because the input size as 320*320 instead of 224*224  for the original VGG model.  \r\n\r\n\r\n**4. What parameters can be experimented for improving the score**\r\n\r\na.  Data augment: train.py has multiple augment options, and one can experiment with mx.io.ImageRecordIter, for example, random rotation, random illumination, and other options. These options can improve the LB scores by generalizing the model. Please refer to MXnet document for how to use them.\r\n\r\nb.  CV: this competition has very small dataset which gives large variance for models. Having multiple runs on multiple different split on train-val dataset can significantly reduce the variance. We provide 5-fold CV, and we have experiment 10-fold CV and other CV tries. Ensembling multiple CV choices can also help smoothing down the variance. Some one has asked if spliting by driver info can help, our experiment shows it is an dangerous choice and it may overfit to some unnecessary features.\r\n\r\nc.  Learning rate change: usually pretrained model needs a large learning rate from new data for quickly converging, so in the train_model.py we have a fine tune learning rate scale parameter `--finetune-lr-scale` for tuning up. There are other parameters like `--lr-factor-epoch` for fine tune the parameters for different epochs.\r\n\r\nd.  Crop in test set: in the vgg directory, one can also find `run.test_10crop.sh` for running a random crop for 10 times and  average on each test image. It can improve LB score from about 0.05 to 0.1, but it takes much longer time. One can modify and get some other crop methods, e.g. 3 crop etc.",
      "votes": null
    },
    {
      "id": "125819",
      "postDate": "07/03/2016 00:11:51",
      "content": "<p>Big thanks to Terry, great work!</p>",
      "rawMarkdown": "Big thanks to Terry, great work!",
      "votes": null
    },
    {
      "id": "125820",
      "postDate": "07/03/2016 00:28:50",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "125838",
      "postDate": "07/03/2016 10:01:33",
      "content": "<p>&quot;Crop in test set: ...  It can improve LB score from about 0.05 to 0.1 &quot;. Thanks for the information. I realize I haven't done this set yet!</p>",
      "rawMarkdown": "\"Crop in test set: ...  It can improve LB score from about 0.05 to 0.1 \". Thanks for the information. I realize I haven't done this set yet!",
      "votes": null
    },
    {
      "id": "126019",
      "postDate": "07/05/2016 15:25:57",
      "content": "<p>Thanks for ur share. \nBut i meet a problem when i try your code, in step 2.b.3), there is always some errors like this: </p>\n\n<p>mxnet.base.MXNetError: [19:24:08] src/io/local_filesys.cc:81: LocalFileSystem.ListDirectory ./rec_224 error: No such file or directory\ncp: cannot stat &#8216;mean.bin&#8217;: No such file or directory</p>\n\n<p>Do u have any idea about this?\nThanks</p>",
      "rawMarkdown": "Thanks for ur share. \r\nBut i meet a problem when i try your code, in step 2.b.3), there is always some errors like this: \r\n\r\nmxnet.base.MXNetError: [19:24:08] src/io/local_filesys.cc:81: LocalFileSystem.ListDirectory ./rec_224 error: No such file or directory\r\ncp: cannot stat ‘mean.bin’: No such file or directory\r\n\r\n\r\nDo u have any idea about this?\r\nThanks",
      "votes": null
    },
    {
      "id": "126021",
      "postDate": "07/05/2016 15:43:46",
      "content": "<p>@zhiqiangZhong - Try changing the default value for --data-dir parameter in train_inception_bn.py to ./ from ./rec_224/ \nHope it helps</p>",
      "rawMarkdown": "zhiqiangZhong - Try changing the default value for --data-dir parameter in train_inception_bn.py to ./ from ./rec_224/ \r\nHope it helps",
      "votes": null
    },
    {
      "id": "126250",
      "postDate": "07/07/2016 05:39:46",
      "content": "<p>@Terry</p>\n\n<p>Thank you for sharing the code!\nI will tweak around with it and hopefully report some good news here soon!</p>\n\n<p>Chris</p>",
      "rawMarkdown": "Terry\r\n\r\nThank you for sharing the code!\r\nI will tweak around with it and hopefully report some good news here soon!\r\n\r\nChris",
      "votes": null
    },
    {
      "id": "126518",
      "postDate": "07/09/2016 04:55:07",
      "content": "<p>Thanks for sharing.  Great work!</p>",
      "rawMarkdown": "Thanks for sharing.  Great work!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 125819,
      "author_name": "phunter",
      "author_url": "",
      "post_date": "07/03/2016 00:11:51",
      "content": "<p>Big thanks to Terry, great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125820,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "07/03/2016 00:28:50",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125838,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/03/2016 10:01:33",
      "content": "<p>&quot;Crop in test set: ...  It can improve LB score from about 0.05 to 0.1 &quot;. Thanks for the information. I realize I haven't done this set yet!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126019,
      "author_name": "zhiqiangzhong",
      "author_url": "",
      "post_date": "07/05/2016 15:25:57",
      "content": "<p>Thanks for ur share. \nBut i meet a problem when i try your code, in step 2.b.3), there is always some errors like this: </p>\n\n<p>mxnet.base.MXNetError: [19:24:08] src/io/local_filesys.cc:81: LocalFileSystem.ListDirectory ./rec_224 error: No such file or directory\ncp: cannot stat &#8216;mean.bin&#8217;: No such file or directory</p>\n\n<p>Do u have any idea about this?\nThanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126021,
      "author_name": "ankdesh",
      "author_url": "",
      "post_date": "07/05/2016 15:43:46",
      "content": "<p>@zhiqiangZhong - Try changing the default value for --data-dir parameter in train_inception_bn.py to ./ from ./rec_224/ \nHope it helps</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126250,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "07/07/2016 05:39:46",
      "content": "<p>@Terry</p>\n\n<p>Thank you for sharing the code!\nI will tweak around with it and hopefully report some good news here soon!</p>\n\n<p>Chris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126518,
      "author_name": "heshameraqi",
      "author_url": "",
      "post_date": "07/09/2016 04:55:07",
      "content": "<p>Thanks for sharing.  Great work!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "125810": "Acknowledgement:  Thanks to Tianqi, Bing, Terry and Hongliang for the code work. MXnet repo https://github.com/dmlc/mxnet please star and fork this great repo. Special thanks to ysfalo for his fine tune blog https://ysfalo.github.io/2016/04/01/mxnet%E4%B9%8Bfine-tune/ \r\n\r\n **1. Github link https://github.com/phunterlau/kaggle_statefarm**\r\n\r\n **2. How to use**\r\n\r\na. Install mxnet by following the official instruction for compiling and installing python packages. Please install necessary CUDA packages before compiling. Please turn on the CUDA option, cuDNN option for fast training (cuDNN v5 is suggested), and the most important, OPENCV option for having ‘im2rec’ binary.\r\n\r\n\r\nb. Run Inception-BN model\r\n\r\n1) Go to inception directory, download the pretrained Inception BN model from dmlc.ml by executing `download.sh`. The pretrained model will appear in `model` directory.\r\n\r\n2) Generate `.rec` format train-val dataset using `im2rec`: put the training dataset into `train` directory, edit `gen_rec_bin.sh` and replace `im2rec_path` variable to the location of  `im2rec` binary. Usually it is in the `bin` directory after compiling MXNet. Run gen_rec_bin.sh and generate 5 cross-validation datasets for 5 batches of training.\r\n\r\n\r\n3) Train the model: run `run.cv_inception_bn.sh` to execute the model training on 5 cross-validation datasets. The code will save each epoch snapshot in `model` directory, and the `mean.bin` for each batch. Please save these snapshots and `mean.bin` for testing. Usually in the first couple of epochs, the training and validation accuracy can increase and reach almost 1.000 after some epochs. By default, this script asks for 30 epochs. Each epoch takes about 150 (Titan X) - 300+ (GTX 960) seconds depending on the video card, so usually 30 epochs for all 5 folds can take around 7 hours on a single decent card with less than 2GB memory. If you have multiple video cards, please use `--gpus` options to use multiple cards, and increase batch size if you want.\r\n\r\n\r\n4) Inference: This is time consuming part because there are about 80k test images. In the inception directory, `run.test.sh` can load the model snapshots for each CV batch, run a center crop of test image, and produce the prediction in the directory, like `submission_test_batch_5_epoch_20.csv`. These output can be merged with other models, or be averaged for a single model submission. The single model inception-BN with the current setup can give 0.59 in the public LB.\r\n\r\n\r\nc. Run VGG model.\r\n\r\n1) Go to VGG directory, download the pretrained VGG models by executing `download.sh`\r\n\r\n2) Generate `.rec` format train-val dataset like you have done for inception BN model. Please note that, VGG and Inception BN use different input size.\r\n\r\n3) Train the model. Run `run.cv.sh` for run a 5 fold 15 epochs VGG model. You might notice `--finetune-lr-scale 5` parameter in the shell script, that is because we have experimented 5x learning rate can quickly produce some good results. Training time for each epoch is about 450 (Titan X) -1000 (GTX 960) seconds, and the total training time for a 5-fold run 15 epoch each can be around 10 hours.\r\n\r\n4) Inference: it is similar as in Inception BN model. With the current setup, the single VGG model can give about 0.56 scores in the public LB.\r\n\r\n\r\nEnsemble different model output: please go to `merge` directory, make two directories `vgg` and `inception`, put the two set of output to the corresponding directories, run `merge.py` and the output is ready to be submitted. The equal weight ensemble of an inception BN model with VGG can reproduce approximately 0.51 LB.\r\n\r\n\r\n**3. How pretrain model works:**\r\n\r\nIn MXnet, we can load a model and keep training with new data. In this problem, we load models from large dataset and continue to train with state farm dataset, that requires that the new data has little impact to the first two layers of the convolutional network, so the network can keep the original weights for the first two CNN layers, and learn from the new dataset for the high level features. In inception model and vgg model’s symbol python file, one can find these convolutional layers having `attr={'lr_mult': '0.00'}` in the layer description. In VGG model, using average pooling instead of two fully connected layer before classifier, because the input size as 320*320 instead of 224*224  for the original VGG model.  \r\n\r\n\r\n**4. What parameters can be experimented for improving the score**\r\n\r\na.  Data augment: train.py has multiple augment options, and one can experiment with mx.io.ImageRecordIter, for example, random rotation, random illumination, and other options. These options can improve the LB scores by generalizing the model. Please refer to MXnet document for how to use them.\r\n\r\nb.  CV: this competition has very small dataset which gives large variance for models. Having multiple runs on multiple different split on train-val dataset can significantly reduce the variance. We provide 5-fold CV, and we have experiment 10-fold CV and other CV tries. Ensembling multiple CV choices can also help smoothing down the variance. Some one has asked if spliting by driver info can help, our experiment shows it is an dangerous choice and it may overfit to some unnecessary features.\r\n\r\nc.  Learning rate change: usually pretrained model needs a large learning rate from new data for quickly converging, so in the train_model.py we have a fine tune learning rate scale parameter `--finetune-lr-scale` for tuning up. There are other parameters like `--lr-factor-epoch` for fine tune the parameters for different epochs.\r\n\r\nd.  Crop in test set: in the vgg directory, one can also find `run.test_10crop.sh` for running a random crop for 10 times and  average on each test image. It can improve LB score from about 0.05 to 0.1, but it takes much longer time. One can modify and get some other crop methods, e.g. 3 crop etc.",
    "125819": "Big thanks to Terry, great work!",
    "125820": "Thanks for sharing!",
    "125838": "\"Crop in test set: ...  It can improve LB score from about 0.05 to 0.1 \". Thanks for the information. I realize I haven't done this set yet!",
    "126019": "Thanks for ur share. \r\nBut i meet a problem when i try your code, in step 2.b.3), there is always some errors like this: \r\n\r\nmxnet.base.MXNetError: [19:24:08] src/io/local_filesys.cc:81: LocalFileSystem.ListDirectory ./rec_224 error: No such file or directory\r\ncp: cannot stat ‘mean.bin’: No such file or directory\r\n\r\n\r\nDo u have any idea about this?\r\nThanks",
    "126021": "zhiqiangZhong - Try changing the default value for --data-dir parameter in train_inception_bn.py to ./ from ./rec_224/ \r\nHope it helps",
    "126250": "Terry\r\n\r\nThank you for sharing the code!\r\nI will tweak around with it and hopefully report some good news here soon!\r\n\r\nChris",
    "126518": "Thanks for sharing.  Great work!"
  },
  "source": "meta"
}