{
  "id": 21778,
  "title": "Pretrained RESNET in Keras ",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/21778",
  "author_name": "",
  "post_date": "2016-06-19T06:36:21.903Z",
  "votes": null,
  "comment_count": 7,
  "views": 3415,
  "content": "<p>Have anyone tried loading prertrained RESNET in Keras and train the model? Appreciate any insights,codes and accuracy on the model.</p>\n\n<p>Thanks\nPatrick</p>",
  "messages": [
    {
      "id": "124483",
      "postDate": "06/19/2016 06:36:21",
      "content": "<p>Have anyone tried loading prertrained RESNET in Keras and train the model? Appreciate any insights,codes and accuracy on the model.</p>\n\n<p>Thanks\nPatrick</p>",
      "rawMarkdown": "Have anyone tried loading prertrained RESNET in Keras and train the model? Appreciate any insights,codes and accuracy on the model.\r\n\r\nThanks\r\nPatrick",
      "votes": null
    },
    {
      "id": "125916",
      "postDate": "07/04/2016 12:54:12",
      "content": "<p>Patrick,</p>\n\n<p>I tried RESNET in Caffe.\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\n<a href=\"http://caffe.berkeleyvision.org/\">http://caffe.berkeleyvision.org/</a></p>\n\n<p>Caffe pretrained model is here:\n<a href=\"https://github.com/KaimingHe/deep-residual-networks\">https://github.com/KaimingHe/deep-residual-networks</a></p>\n\n<p>To my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.</p>\n\n<p>With ResNet-50, single best model got me LB ~0.32</p>\n\n<p>Good Luck,\nChris</p>",
      "rawMarkdown": "Patrick,\r\n\r\nI tried RESNET in Caffe.\r\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\r\nhttp://caffe.berkeleyvision.org/\r\n\r\nCaffe pretrained model is here:\r\nhttps://github.com/KaimingHe/deep-residual-networks\r\n\r\nTo my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.\r\n\r\nWith ResNet-50, single best model got me LB ~0.32\r\n\r\nGood Luck,\r\nChris",
      "votes": null
    },
    {
      "id": "125952",
      "postDate": "07/04/2016 23:09:52",
      "content": "<p>@ChrisJung\ncan you gives details on how you get LB 0.34?\ne.g solver network and prototxt file parameters, like learning rate, augmentation how many epoch,training samples?\nI have been working with resNet for weeks but cannot get good results. </p>\n\n<p>[quote=ChrisJung;125916]</p>\n\n<p>Patrick,</p>\n\n<p>I tried RESNET in Caffe.\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\n<a href=\"http://caffe.berkeleyvision.org/\">http://caffe.berkeleyvision.org/</a></p>\n\n<p>Caffe pretrained model is here:\n<a href=\"https://github.com/KaimingHe/deep-residual-networks\">https://github.com/KaimingHe/deep-residual-networks</a></p>\n\n<p>To my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.</p>\n\n<p>With ResNet-50, single best model got me LB ~0.32</p>\n\n<p>Good Luck,\nChris</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "ChrisJung\r\ncan you gives details on how you get LB 0.34?\r\ne.g solver network and prototxt file parameters, like learning rate, augmentation how many epoch,training samples?\r\nI have been working with resNet for weeks but cannot get good results. \r\n\r\n\r\n[quote=ChrisJung;125916]\r\n\r\nPatrick,\r\n\r\nI tried RESNET in Caffe.\r\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\r\nhttp://caffe.berkeleyvision.org/\r\n\r\nCaffe pretrained model is here:\r\nhttps://github.com/KaimingHe/deep-residual-networks\r\n\r\nTo my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.\r\n\r\nWith ResNet-50, single best model got me LB ~0.32\r\n\r\nGood Luck,\r\nChris\r\n\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "126075",
      "postDate": "07/06/2016 01:25:52",
      "content": "<p>@Heng Cherkeng</p>\n\n<p>Below are some details:</p>\n\n<ol>\n<li><p>model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.</p></li>\n<li><p>solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)</p></li>\n<li><p>No data augmentation at all. Divided my train/val dataset randomly on driver_id.</p></li>\n</ol>\n\n<p>Personally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.</p>\n\n<p>Attached is the learning curve graph.\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\nIt seems there is still some overfitting by ResNet-50. \nHope to tackle these with some cool tricks. :)</p>\n\n<p>Hope you share me your working method too!</p>\n\n<p>Chris</p>",
      "rawMarkdown": "Heng Cherkeng\r\n\r\nBelow are some details:\r\n\r\n1. model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.\r\n\r\n2. solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)\r\n\r\n3. No data augmentation at all. Divided my train/val dataset randomly on driver_id.\r\n\r\nPersonally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\r\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.\r\n\r\nAttached is the learning curve graph.\r\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\r\nIt seems there is still some overfitting by ResNet-50. \r\nHope to tackle these with some cool tricks. :)\r\n\r\nHope you share me your working method too!\r\n\r\nChris",
      "votes": null
    },
    {
      "id": "126078",
      "postDate": "07/06/2016 01:33:52",
      "content": "<p>@ChrisJung\nThank you very much! this is very helpful. Now I try resNet again. Will keep you update when I have results.</p>\n\n<p>Meanwhile, I am also working on vgg16_CAM and googlenet-CAM. For these results, please refer to:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/126077#post126077\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/126077#post126077</a></p>\n\n<p>[quote=ChrisJung;126075]</p>\n\n<p>@Heng Cherkeng</p>\n\n<p>Below are some details:</p>\n\n<ol>\n<li><p>model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.</p></li>\n<li><p>solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)</p></li>\n<li><p>No data augmentation at all. Divided my train/val dataset randomly on driver_id.</p></li>\n</ol>\n\n<p>Personally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.</p>\n\n<p>Attached is the learning curve graph.\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\nIt seems there is still some overfitting by ResNet-50. \nHope to tackle these with some cool tricks. :)</p>\n\n<p>Hope you share me your working method too!</p>\n\n<p>Chris</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "ChrisJung\r\nThank you very much! this is very helpful. Now I try resNet again. Will keep you update when I have results.\r\n\r\nMeanwhile, I am also working on vgg16_CAM and googlenet-CAM. For these results, please refer to:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/126077#post126077\r\n\r\n[quote=ChrisJung;126075]\r\n\r\n@Heng Cherkeng\r\n\r\nBelow are some details:\r\n\r\n1. model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.\r\n\r\n2. solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)\r\n\r\n3. No data augmentation at all. Divided my train/val dataset randomly on driver_id.\r\n\r\nPersonally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\r\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.\r\n\r\nAttached is the learning curve graph.\r\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\r\nIt seems there is still some overfitting by ResNet-50. \r\nHope to tackle these with some cool tricks. :)\r\n\r\nHope you share me your working method too!\r\n\r\nChris\r\n\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "126112",
      "postDate": "07/06/2016 11:14:42",
      "content": "<p>@Heng CherKeng</p>\n\n<p>No problem :)\nI saw your posts on the Forum for ideas and good references.\nI will also update you my result after I try out your ideas.</p>\n\n<p>Thanks,\nChris</p>",
      "rawMarkdown": "Heng CherKeng\r\n\r\nNo problem :)\r\nI saw your posts on the Forum for ideas and good references.\r\nI will also update you my result after I try out your ideas.\r\n\r\nThanks,\r\nChris",
      "votes": null
    },
    {
      "id": "126123",
      "postDate": "07/06/2016 15:07:36",
      "content": "<p>Can you please send us your training and testing time for rasnet (and your GPU type)? </p>",
      "rawMarkdown": "Can you please send us your training and testing time for rasnet (and your GPU type)?",
      "votes": null
    },
    {
      "id": "126231",
      "postDate": "07/07/2016 04:28:07",
      "content": "<p>@Ehsan</p>\n\n<p>My device status is as below:\nOS: Ubuntu 14.04 Server\nCPU: Intel Core i7-6700 @ 3.4GHz (8 cores)\nMemory: 16GB\nGPU: GeForce GTX970 with 4GB Memory</p>\n\n<p>With CuDNN v5 installed Caffe, training time took about 9 hours.\nWhen you use LMDB dataset, it's usually faster.\nImageData layer would be much slower (~12hours)</p>\n\n<p>Testing time(getting predictions on test data) takes around &lt; 2 hours for me.</p>\n\n<p>[quote=Ehsan;126123]</p>\n\n<p>Can you please send us your training and testing time for rasnet (and your GPU type)? </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "Ehsan\r\n\r\nMy device status is as below:\r\nOS: Ubuntu 14.04 Server\r\nCPU: Intel Core i7-6700 @ 3.4GHz (8 cores)\r\nMemory: 16GB\r\nGPU: GeForce GTX970 with 4GB Memory\r\n\r\nWith CuDNN v5 installed Caffe, training time took about 9 hours.\r\nWhen you use LMDB dataset, it's usually faster.\r\nImageData layer would be much slower (~12hours)\r\n\r\nTesting time(getting predictions on test data) takes around < 2 hours for me.\r\n\r\n[quote=Ehsan;126123]\r\n\r\nCan you please send us your training and testing time for rasnet (and your GPU type)? \r\n\r\n[/quote]",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 125916,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "07/04/2016 12:54:12",
      "content": "<p>Patrick,</p>\n\n<p>I tried RESNET in Caffe.\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\n<a href=\"http://caffe.berkeleyvision.org/\">http://caffe.berkeleyvision.org/</a></p>\n\n<p>Caffe pretrained model is here:\n<a href=\"https://github.com/KaimingHe/deep-residual-networks\">https://github.com/KaimingHe/deep-residual-networks</a></p>\n\n<p>To my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.</p>\n\n<p>With ResNet-50, single best model got me LB ~0.32</p>\n\n<p>Good Luck,\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125952,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/04/2016 23:09:52",
      "content": "<p>@ChrisJung\ncan you gives details on how you get LB 0.34?\ne.g solver network and prototxt file parameters, like learning rate, augmentation how many epoch,training samples?\nI have been working with resNet for weeks but cannot get good results. </p>\n\n<p>[quote=ChrisJung;125916]</p>\n\n<p>Patrick,</p>\n\n<p>I tried RESNET in Caffe.\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\n<a href=\"http://caffe.berkeleyvision.org/\">http://caffe.berkeleyvision.org/</a></p>\n\n<p>Caffe pretrained model is here:\n<a href=\"https://github.com/KaimingHe/deep-residual-networks\">https://github.com/KaimingHe/deep-residual-networks</a></p>\n\n<p>To my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.</p>\n\n<p>With ResNet-50, single best model got me LB ~0.32</p>\n\n<p>Good Luck,\nChris</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126075,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "07/06/2016 01:25:52",
      "content": "<p>@Heng Cherkeng</p>\n\n<p>Below are some details:</p>\n\n<ol>\n<li><p>model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.</p></li>\n<li><p>solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)</p></li>\n<li><p>No data augmentation at all. Divided my train/val dataset randomly on driver_id.</p></li>\n</ol>\n\n<p>Personally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.</p>\n\n<p>Attached is the learning curve graph.\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\nIt seems there is still some overfitting by ResNet-50. \nHope to tackle these with some cool tricks. :)</p>\n\n<p>Hope you share me your working method too!</p>\n\n<p>Chris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126078,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "07/06/2016 01:33:52",
      "content": "<p>@ChrisJung\nThank you very much! this is very helpful. Now I try resNet again. Will keep you update when I have results.</p>\n\n<p>Meanwhile, I am also working on vgg16_CAM and googlenet-CAM. For these results, please refer to:\n<a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/126077#post126077\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/126077#post126077</a></p>\n\n<p>[quote=ChrisJung;126075]</p>\n\n<p>@Heng Cherkeng</p>\n\n<p>Below are some details:</p>\n\n<ol>\n<li><p>model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.</p></li>\n<li><p>solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)</p></li>\n<li><p>No data augmentation at all. Divided my train/val dataset randomly on driver_id.</p></li>\n</ol>\n\n<p>Personally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.</p>\n\n<p>Attached is the learning curve graph.\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\nIt seems there is still some overfitting by ResNet-50. \nHope to tackle these with some cool tricks. :)</p>\n\n<p>Hope you share me your working method too!</p>\n\n<p>Chris</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126112,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "07/06/2016 11:14:42",
      "content": "<p>@Heng CherKeng</p>\n\n<p>No problem :)\nI saw your posts on the Forum for ideas and good references.\nI will also update you my result after I try out your ideas.</p>\n\n<p>Thanks,\nChris</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126123,
      "author_name": "ehsanma",
      "author_url": "",
      "post_date": "07/06/2016 15:07:36",
      "content": "<p>Can you please send us your training and testing time for rasnet (and your GPU type)? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126231,
      "author_name": "kweonwooj",
      "author_url": "",
      "post_date": "07/07/2016 04:28:07",
      "content": "<p>@Ehsan</p>\n\n<p>My device status is as below:\nOS: Ubuntu 14.04 Server\nCPU: Intel Core i7-6700 @ 3.4GHz (8 cores)\nMemory: 16GB\nGPU: GeForce GTX970 with 4GB Memory</p>\n\n<p>With CuDNN v5 installed Caffe, training time took about 9 hours.\nWhen you use LMDB dataset, it's usually faster.\nImageData layer would be much slower (~12hours)</p>\n\n<p>Testing time(getting predictions on test data) takes around &lt; 2 hours for me.</p>\n\n<p>[quote=Ehsan;126123]</p>\n\n<p>Can you please send us your training and testing time for rasnet (and your GPU type)? </p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "124483": "Have anyone tried loading prertrained RESNET in Keras and train the model? Appreciate any insights,codes and accuracy on the model.\r\n\r\nThanks\r\nPatrick",
    "125916": "Patrick,\r\n\r\nI tried RESNET in Caffe.\r\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\r\nhttp://caffe.berkeleyvision.org/\r\n\r\nCaffe pretrained model is here:\r\nhttps://github.com/KaimingHe/deep-residual-networks\r\n\r\nTo my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.\r\n\r\nWith ResNet-50, single best model got me LB ~0.32\r\n\r\nGood Luck,\r\nChris",
    "125952": "ChrisJung\r\ncan you gives details on how you get LB 0.34?\r\ne.g solver network and prototxt file parameters, like learning rate, augmentation how many epoch,training samples?\r\nI have been working with resNet for weeks but cannot get good results. \r\n\r\n\r\n[quote=ChrisJung;125916]\r\n\r\nPatrick,\r\n\r\nI tried RESNET in Caffe.\r\nTook few days to install and run Caffe on my Ubuntu Server, but definitely worth it.\r\nhttp://caffe.berkeleyvision.org/\r\n\r\nCaffe pretrained model is here:\r\nhttps://github.com/KaimingHe/deep-residual-networks\r\n\r\nTo my knowledge, there is no pretrained RESNET model for KERAS. That's why I dug into Caffe myself.\r\n\r\nWith ResNet-50, single best model got me LB ~0.32\r\n\r\nGood Luck,\r\nChris\r\n\r\n\r\n[/quote]",
    "126075": "Heng Cherkeng\r\n\r\nBelow are some details:\r\n\r\n1. model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.\r\n\r\n2. solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)\r\n\r\n3. No data augmentation at all. Divided my train/val dataset randomly on driver_id.\r\n\r\nPersonally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\r\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.\r\n\r\nAttached is the learning curve graph.\r\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\r\nIt seems there is still some overfitting by ResNet-50. \r\nHope to tackle these with some cool tricks. :)\r\n\r\nHope you share me your working method too!\r\n\r\nChris",
    "126078": "ChrisJung\r\nThank you very much! this is very helpful. Now I try resNet again. Will keep you update when I have results.\r\n\r\nMeanwhile, I am also working on vgg16_CAM and googlenet-CAM. For these results, please refer to:\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/21994/heat-map-of-cnn-output/126077#post126077\r\n\r\n[quote=ChrisJung;126075]\r\n\r\n@Heng Cherkeng\r\n\r\nBelow are some details:\r\n\r\n1. model.prototxt : Only modified data layer (into LMDB) and final fc layer (from fc1000 to fc10). Also added loss layer, accuracy layer for monitoring purpose.\r\n\r\n2. solver.prototxt : I tried the classic SGD method. Total of 20 epochs. learning rate of 0.001 for first 10 epochs and 0.0001 for next 10. (hence gamma=0.1) Momentum=0.9 and weight_decay = 0.0005 as default. ( So far, this is my best hyperparameter for single model)\r\n\r\n3. No data augmentation at all. Divided my train/val dataset randomly on driver_id.\r\n\r\nPersonally, I tried ResNet-50, 101, 152 and ResNet-50 seems to perform best.\r\nPerhaps, I haven't tried out all the hyperparameters or maybe ResNet-100+ is too deep for this competition.\r\n\r\nAttached is the learning curve graph.\r\nI used gamma = 0.5 per 5 epochs (I'm using batch_size=4 so ~5000 iterations = 1 epoch)\r\nIt seems there is still some overfitting by ResNet-50. \r\nHope to tackle these with some cool tricks. :)\r\n\r\nHope you share me your working method too!\r\n\r\nChris\r\n\r\n\r\n[/quote]",
    "126112": "Heng CherKeng\r\n\r\nNo problem :)\r\nI saw your posts on the Forum for ideas and good references.\r\nI will also update you my result after I try out your ideas.\r\n\r\nThanks,\r\nChris",
    "126123": "Can you please send us your training and testing time for rasnet (and your GPU type)?",
    "126231": "Ehsan\r\n\r\nMy device status is as below:\r\nOS: Ubuntu 14.04 Server\r\nCPU: Intel Core i7-6700 @ 3.4GHz (8 cores)\r\nMemory: 16GB\r\nGPU: GeForce GTX970 with 4GB Memory\r\n\r\nWith CuDNN v5 installed Caffe, training time took about 9 hours.\r\nWhen you use LMDB dataset, it's usually faster.\r\nImageData layer would be much slower (~12hours)\r\n\r\nTesting time(getting predictions on test data) takes around < 2 hours for me.\r\n\r\n[quote=Ehsan;126123]\r\n\r\nCan you please send us your training and testing time for rasnet (and your GPU type)? \r\n\r\n[/quote]"
  },
  "source": "meta"
}