{
  "id": 20674,
  "title": "Training on CPU (no GPU)",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20674",
  "author_name": "",
  "post_date": "2016-05-03T14:43:31.703Z",
  "votes": 1,
  "comment_count": 8,
  "views": 1195,
  "content": "<p>Hi,</p>\n\n<p>I am wondering if anybody has been able to make it to the top 50 without using a GPU, running their model on a CPU. If there is anyone, I would like to know how much time their script takes. </p>\n\n<p>I do not have a GPU, and my Keras based model (a modified version of <a href=\"https://www.kaggle.com/zfturbo/state-farm-distracted-driver-detection/keras-sample\">ZFTurbo's script</a>) takes around 1.5 days for a 10-fold CV to produce a log_loss of ~1.25 on all the folds combined. After blending the results of several runs with a slight change of parameters, I am able to get a log_loss of ~1.08 on the LB. </p>\n\n<p>My model has the following architecture: Conv-Conv-MaxPool-Dropout-Conv-Conv-MaxPool-Dropout-Dense-Dropout-Dense-Dropout-Output. First 2 convolution layers compute 32 kernels and the second set of convolution layers compute 64 kernels of size 3x3. The Dense layers have 128 units each. All the activations, except the last layer, are ReLU. I am training the n/w with 24x32 grey scale images using the adadelta optimizer.</p>\n\n<p>I would appreciate any tips on making my model run faster on a CPU. Thanks!</p>",
  "messages": [
    {
      "id": "118402",
      "postDate": "05/03/2016 14:43:31",
      "content": "<p>Hi,</p>\n\n<p>I am wondering if anybody has been able to make it to the top 50 without using a GPU, running their model on a CPU. If there is anyone, I would like to know how much time their script takes. </p>\n\n<p>I do not have a GPU, and my Keras based model (a modified version of <a href=\"https://www.kaggle.com/zfturbo/state-farm-distracted-driver-detection/keras-sample\">ZFTurbo's script</a>) takes around 1.5 days for a 10-fold CV to produce a log_loss of ~1.25 on all the folds combined. After blending the results of several runs with a slight change of parameters, I am able to get a log_loss of ~1.08 on the LB. </p>\n\n<p>My model has the following architecture: Conv-Conv-MaxPool-Dropout-Conv-Conv-MaxPool-Dropout-Dense-Dropout-Dense-Dropout-Output. First 2 convolution layers compute 32 kernels and the second set of convolution layers compute 64 kernels of size 3x3. The Dense layers have 128 units each. All the activations, except the last layer, are ReLU. I am training the n/w with 24x32 grey scale images using the adadelta optimizer.</p>\n\n<p>I would appreciate any tips on making my model run faster on a CPU. Thanks!</p>",
      "rawMarkdown": "Hi,\r\n\r\nI am wondering if anybody has been able to make it to the top 50 without using a GPU, running their model on a CPU. If there is anyone, I would like to know how much time their script takes. \r\n\r\nI do not have a GPU, and my Keras based model (a modified version of [ZFTurbo's script][1]) takes around 1.5 days for a 10-fold CV to produce a log_loss of ~1.25 on all the folds combined. After blending the results of several runs with a slight change of parameters, I am able to get a log_loss of ~1.08 on the LB. \r\n\r\nMy model has the following architecture: Conv-Conv-MaxPool-Dropout-Conv-Conv-MaxPool-Dropout-Dense-Dropout-Dense-Dropout-Output. First 2 convolution layers compute 32 kernels and the second set of convolution layers compute 64 kernels of size 3x3. The Dense layers have 128 units each. All the activations, except the last layer, are ReLU. I am training the n/w with 24x32 grey scale images using the adadelta optimizer.\r\n\r\nI would appreciate any tips on making my model run faster on a CPU. Thanks!\r\n\r\n  [1]: https://www.kaggle.com/zfturbo/state-farm-distracted-driver-detection/keras-sample",
      "votes": null
    },
    {
      "id": "118427",
      "postDate": "05/03/2016 17:01:42",
      "content": "<p>I don't have a GPU either. I have a similar model to yours, takes about an hour to train on a CPU - but then I run it 13-fold, so it takes a little over half a day to train - I'm guessing your 1.5 days includes the multiple k-fold approach?</p>\n\n<p>My best architecture is very similar to yours. Also using adadelta - I tried all the optimisers in Keras, and this one seemed to get best CV in shortest time for the current setup. I've been using leaky ReLU, which did seem to help a little. Of course I cannot explore all options due to training time, so may be missing something. Also, I am always concerned that just doubling the number of filters, or doubling size of base layer and having an extra pooling layer, would be very effective but totally out of the question for training time.</p>",
      "rawMarkdown": "I don't have a GPU either. I have a similar model to yours, takes about an hour to train on a CPU - but then I run it 13-fold, so it takes a little over half a day to train - I'm guessing your 1.5 days includes the multiple k-fold approach?\r\n\r\nMy best architecture is very similar to yours. Also using adadelta - I tried all the optimisers in Keras, and this one seemed to get best CV in shortest time for the current setup. I've been using leaky ReLU, which did seem to help a little. Of course I cannot explore all options due to training time, so may be missing something. Also, I am always concerned that just doubling the number of filters, or doubling size of base layer and having an extra pooling layer, would be very effective but totally out of the question for training time.",
      "votes": null
    },
    {
      "id": "118516",
      "postDate": "05/04/2016 04:12:49",
      "content": "<p>Thanks for the response Neil. Actually it's taking more than a day for a single run of a 10-fold CV. I am using early stopping with a patience level of 5. Each epoch is taking around 11 minutes and on an average each fold runs for around 20 epochs. That amounts to a total of around 37 hours.</p>\n\n<p>What I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. </p>\n\n<p>Now, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!</p>",
      "rawMarkdown": "Thanks for the response Neil. Actually it's taking more than a day for a single run of a 10-fold CV. I am using early stopping with a patience level of 5. Each epoch is taking around 11 minutes and on an average each fold runs for around 20 epochs. That amounts to a total of around 37 hours.\r\n\r\nWhat I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. \r\n\r\nNow, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!",
      "votes": null
    },
    {
      "id": "118538",
      "postDate": "05/04/2016 07:27:16",
      "content": "<p>[quote=msram;118516]</p>\n\n<p>What I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. </p>\n\n<p>[/quote]</p>\n\n<p>I've been using that too. I have been stopping runs that look either too similar to previous ones, or clearly far worse (worse training and cv scores).</p>\n\n<p>[quote=msram;118516]</p>\n\n<p>Now, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!</p>\n\n<p>[/quote]</p>\n\n<p>A &quot;single fold&quot; is just a normal hold-out CV. It is a standard approach to pick meta-params and/or models, and I've worked with a single hold-out set with NNs in other competitions and got top 10% scores (Right Whale Detection and Africa Soil Quality). However, I think it is a risk with this dataset, because CV for each driver varies so much. I have seen individual fold CV scores varying by 0.2 where the LB score is essentially the same once the models are combined.</p>\n\n<p>I'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.</p>\n\n<p>I think you will get less accurate measures of generalisation locally, but you will be able to build a meta-model faster, and get some feedback about that from the public test set. I cannot predict whether that's a better strategy when the limitation is with CPU time. I think it is a risk, but maybe one worth taking. However, I won't be doing that. I think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.</p>",
      "rawMarkdown": "[quote=msram;118516]\r\n\r\nWhat I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. \r\n\r\n[/quote]\r\n\r\nI've been using that too. I have been stopping runs that look either too similar to previous ones, or clearly far worse (worse training and cv scores).\r\n\r\n[quote=msram;118516]\r\n\r\nNow, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!\r\n\r\n[/quote]\r\n\r\nA \"single fold\" is just a normal hold-out CV. It is a standard approach to pick meta-params and/or models, and I've worked with a single hold-out set with NNs in other competitions and got top 10% scores (Right Whale Detection and Africa Soil Quality). However, I think it is a risk with this dataset, because CV for each driver varies so much. I have seen individual fold CV scores varying by 0.2 where the LB score is essentially the same once the models are combined.\r\n\r\nI'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.\r\n\r\nI think you will get less accurate measures of generalisation locally, but you will be able to build a meta-model faster, and get some feedback about that from the public test set. I cannot predict whether that's a better strategy when the limitation is with CPU time. I think it is a risk, but maybe one worth taking. However, I won't be doing that. I think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.",
      "votes": null
    },
    {
      "id": "118556",
      "postDate": "05/04/2016 08:58:33",
      "content": "<p>[quote=Neil Slater;118538]</p>\n\n<p>I'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.</p>\n\n<p>[/quote]</p>\n\n<p>I agree with you. We just witnessed disastrous differences between the public LB ranks and the private ones in the Santander competition. I wouldn't want to overfit to the public LB either.</p>\n\n<p>[quote=Neil Slater;118538]</p>\n\n<p>I think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.</p>\n\n<p>[/quote]</p>\n\n<p>Have you used data augmentation before? If so, how well did it do? May be I will also try that.</p>",
      "rawMarkdown": "[quote=Neil Slater;118538]\r\n\r\nI'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.\r\n\r\n[/quote]\r\n\r\nI agree with you. We just witnessed disastrous differences between the public LB ranks and the private ones in the Santander competition. I wouldn't want to overfit to the public LB either.\r\n\r\n[quote=Neil Slater;118538]\r\n\r\nI think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.\r\n\r\n[/quote]\r\n\r\nHave you used data augmentation before? If so, how well did it do? May be I will also try that.",
      "votes": null
    },
    {
      "id": "118558",
      "postDate": "05/04/2016 09:05:53",
      "content": "<p>you can check the data augmentation suggested by hoyt in the <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12842/starter-code-for-leaderboard-score-of-0-46\">diabetic retinopathy</a> competition. \nit might be a good starting point.</p>",
      "rawMarkdown": "you can check the data augmentation suggested by hoyt in the [diabetic retinopathy][1] competition. \r\nit might be a good starting point.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12842/starter-code-for-leaderboard-score-of-0-46",
      "votes": null
    },
    {
      "id": "118577",
      "postDate": "05/04/2016 10:51:35",
      "content": "<p>@Nathaniel Shimoni: Thanks for the link.</p>",
      "rawMarkdown": "Nathaniel Shimoni: Thanks for the link.",
      "votes": null
    },
    {
      "id": "118582",
      "postDate": "05/04/2016 11:20:21",
      "content": "<p>[quote=msram;118556]</p>\n\n<p>Have you used data augmentation before? If so, how well did it do? May be I will also try that.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, it did very well for me in the Right Whale Detection competition. In that competition I was using the FANN library (in Ruby), which is not competitive with CNNs. But with data augmentation plus ensembling the end result was pretty good.</p>\n\n<p>Like feature engineering, there is an element of domain knowledge that helps you guess what kind of changes are helpful. I'd guess a variety of small affine transformations are going to be helpful, and also  brightness/contrast.</p>\n\n<p>The alternative to data augmentation, would be normalisation and/or automated cropping - if you can reliably extract the driver's head into a fixed pose form each image, then you have much less need to generate variations of the existing poses.</p>",
      "rawMarkdown": "[quote=msram;118556]\r\n\r\nHave you used data augmentation before? If so, how well did it do? May be I will also try that.\r\n\r\n[/quote]\r\n\r\nYes, it did very well for me in the Right Whale Detection competition. In that competition I was using the FANN library (in Ruby), which is not competitive with CNNs. But with data augmentation plus ensembling the end result was pretty good.\r\n\r\nLike feature engineering, there is an element of domain knowledge that helps you guess what kind of changes are helpful. I'd guess a variety of small affine transformations are going to be helpful, and also  brightness/contrast.\r\n\r\nThe alternative to data augmentation, would be normalisation and/or automated cropping - if you can reliably extract the driver's head into a fixed pose form each image, then you have much less need to generate variations of the existing poses.",
      "votes": null
    },
    {
      "id": "118667",
      "postDate": "05/04/2016 16:57:53",
      "content": "<p>@Neil Slater: I see. I will try these things one by one. Thanks for the detailed response.</p>",
      "rawMarkdown": "Neil Slater: I see. I will try these things one by one. Thanks for the detailed response.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 118427,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "05/03/2016 17:01:42",
      "content": "<p>I don't have a GPU either. I have a similar model to yours, takes about an hour to train on a CPU - but then I run it 13-fold, so it takes a little over half a day to train - I'm guessing your 1.5 days includes the multiple k-fold approach?</p>\n\n<p>My best architecture is very similar to yours. Also using adadelta - I tried all the optimisers in Keras, and this one seemed to get best CV in shortest time for the current setup. I've been using leaky ReLU, which did seem to help a little. Of course I cannot explore all options due to training time, so may be missing something. Also, I am always concerned that just doubling the number of filters, or doubling size of base layer and having an extra pooling layer, would be very effective but totally out of the question for training time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118516,
      "author_name": "techieram",
      "author_url": "",
      "post_date": "05/04/2016 04:12:49",
      "content": "<p>Thanks for the response Neil. Actually it's taking more than a day for a single run of a 10-fold CV. I am using early stopping with a patience level of 5. Each epoch is taking around 11 minutes and on an average each fold runs for around 20 epochs. That amounts to a total of around 37 hours.</p>\n\n<p>What I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. </p>\n\n<p>Now, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118538,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "05/04/2016 07:27:16",
      "content": "<p>[quote=msram;118516]</p>\n\n<p>What I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. </p>\n\n<p>[/quote]</p>\n\n<p>I've been using that too. I have been stopping runs that look either too similar to previous ones, or clearly far worse (worse training and cv scores).</p>\n\n<p>[quote=msram;118516]</p>\n\n<p>Now, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!</p>\n\n<p>[/quote]</p>\n\n<p>A &quot;single fold&quot; is just a normal hold-out CV. It is a standard approach to pick meta-params and/or models, and I've worked with a single hold-out set with NNs in other competitions and got top 10% scores (Right Whale Detection and Africa Soil Quality). However, I think it is a risk with this dataset, because CV for each driver varies so much. I have seen individual fold CV scores varying by 0.2 where the LB score is essentially the same once the models are combined.</p>\n\n<p>I'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.</p>\n\n<p>I think you will get less accurate measures of generalisation locally, but you will be able to build a meta-model faster, and get some feedback about that from the public test set. I cannot predict whether that's a better strategy when the limitation is with CPU time. I think it is a risk, but maybe one worth taking. However, I won't be doing that. I think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118556,
      "author_name": "techieram",
      "author_url": "",
      "post_date": "05/04/2016 08:58:33",
      "content": "<p>[quote=Neil Slater;118538]</p>\n\n<p>I'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.</p>\n\n<p>[/quote]</p>\n\n<p>I agree with you. We just witnessed disastrous differences between the public LB ranks and the private ones in the Santander competition. I wouldn't want to overfit to the public LB either.</p>\n\n<p>[quote=Neil Slater;118538]</p>\n\n<p>I think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.</p>\n\n<p>[/quote]</p>\n\n<p>Have you used data augmentation before? If so, how well did it do? May be I will also try that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118558,
      "author_name": "interesting",
      "author_url": "",
      "post_date": "05/04/2016 09:05:53",
      "content": "<p>you can check the data augmentation suggested by hoyt in the <a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12842/starter-code-for-leaderboard-score-of-0-46\">diabetic retinopathy</a> competition. \nit might be a good starting point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118577,
      "author_name": "techieram",
      "author_url": "",
      "post_date": "05/04/2016 10:51:35",
      "content": "<p>@Nathaniel Shimoni: Thanks for the link.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118582,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "05/04/2016 11:20:21",
      "content": "<p>[quote=msram;118556]</p>\n\n<p>Have you used data augmentation before? If so, how well did it do? May be I will also try that.</p>\n\n<p>[/quote]</p>\n\n<p>Yes, it did very well for me in the Right Whale Detection competition. In that competition I was using the FANN library (in Ruby), which is not competitive with CNNs. But with data augmentation plus ensembling the end result was pretty good.</p>\n\n<p>Like feature engineering, there is an element of domain knowledge that helps you guess what kind of changes are helpful. I'd guess a variety of small affine transformations are going to be helpful, and also  brightness/contrast.</p>\n\n<p>The alternative to data augmentation, would be normalisation and/or automated cropping - if you can reliably extract the driver's head into a fixed pose form each image, then you have much less need to generate variations of the existing poses.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 118667,
      "author_name": "techieram",
      "author_url": "",
      "post_date": "05/04/2016 16:57:53",
      "content": "<p>@Neil Slater: I see. I will try these things one by one. Thanks for the detailed response.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "118402": "Hi,\r\n\r\nI am wondering if anybody has been able to make it to the top 50 without using a GPU, running their model on a CPU. If there is anyone, I would like to know how much time their script takes. \r\n\r\nI do not have a GPU, and my Keras based model (a modified version of [ZFTurbo's script][1]) takes around 1.5 days for a 10-fold CV to produce a log_loss of ~1.25 on all the folds combined. After blending the results of several runs with a slight change of parameters, I am able to get a log_loss of ~1.08 on the LB. \r\n\r\nMy model has the following architecture: Conv-Conv-MaxPool-Dropout-Conv-Conv-MaxPool-Dropout-Dense-Dropout-Dense-Dropout-Output. First 2 convolution layers compute 32 kernels and the second set of convolution layers compute 64 kernels of size 3x3. The Dense layers have 128 units each. All the activations, except the last layer, are ReLU. I am training the n/w with 24x32 grey scale images using the adadelta optimizer.\r\n\r\nI would appreciate any tips on making my model run faster on a CPU. Thanks!\r\n\r\n  [1]: https://www.kaggle.com/zfturbo/state-farm-distracted-driver-detection/keras-sample",
    "118427": "I don't have a GPU either. I have a similar model to yours, takes about an hour to train on a CPU - but then I run it 13-fold, so it takes a little over half a day to train - I'm guessing your 1.5 days includes the multiple k-fold approach?\r\n\r\nMy best architecture is very similar to yours. Also using adadelta - I tried all the optimisers in Keras, and this one seemed to get best CV in shortest time for the current setup. I've been using leaky ReLU, which did seem to help a little. Of course I cannot explore all options due to training time, so may be missing something. Also, I am always concerned that just doubling the number of filters, or doubling size of base layer and having an extra pooling layer, would be very effective but totally out of the question for training time.",
    "118516": "Thanks for the response Neil. Actually it's taking more than a day for a single run of a 10-fold CV. I am using early stopping with a patience level of 5. Each epoch is taking around 11 minutes and on an average each fold runs for around 20 epochs. That amounts to a total of around 37 hours.\r\n\r\nWhat I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. \r\n\r\nNow, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!",
    "118538": "[quote=msram;118516]\r\n\r\nWhat I did was to try different values of hyper parameters on just the first fold and come to an opinion based on how it does on that fold and kill the process and try a different value on the same fold and so on. As a result, my current model does well on the first fold, but I am sure there are other models of similar complexity, which do better on the other folds and overall. \r\n\r\n[/quote]\r\n\r\nI've been using that too. I have been stopping runs that look either too similar to previous ones, or clearly far worse (worse training and cv scores).\r\n\r\n[quote=msram;118516]\r\n\r\nNow, I am tempted to think of an ensembling method that combines models each of which performs well on a single fold, to yield a meta-model that has a good generalizes well to unseen data. Before I invest my time on this problem, I would like to know if it make sense at all or there are any obvious flaws with this line of thinking. Any expert advice on this will be highly appreciated. Thanks!\r\n\r\n[/quote]\r\n\r\nA \"single fold\" is just a normal hold-out CV. It is a standard approach to pick meta-params and/or models, and I've worked with a single hold-out set with NNs in other competitions and got top 10% scores (Right Whale Detection and Africa Soil Quality). However, I think it is a risk with this dataset, because CV for each driver varies so much. I have seen individual fold CV scores varying by 0.2 where the LB score is essentially the same once the models are combined.\r\n\r\nI'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.\r\n\r\nI think you will get less accurate measures of generalisation locally, but you will be able to build a meta-model faster, and get some feedback about that from the public test set. I cannot predict whether that's a better strategy when the limitation is with CPU time. I think it is a risk, but maybe one worth taking. However, I won't be doing that. I think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.",
    "118556": "[quote=Neil Slater;118538]\r\n\r\nI'm not too sure whether you are proposing to vary the CV set between each attempt to tune params for an architecture? If so, then yes I think that will work well for generalisation, but you will lose the ability to assess the meta-model independently of the LB score. Which means you end up tuning to the public LB, and a risk of over-fitting that compared to the private LB.\r\n\r\n[/quote]\r\n\r\nI agree with you. We just witnessed disastrous differences between the public LB ranks and the private ones in the Santander competition. I wouldn't want to overfit to the public LB either.\r\n\r\n[quote=Neil Slater;118538]\r\n\r\nI think my next thing to look at is data augmentation - it will increase training time, but I am gambling that the extra time will be worth it with better generalisation to other car interiors and camera angles.\r\n\r\n[/quote]\r\n\r\nHave you used data augmentation before? If so, how well did it do? May be I will also try that.",
    "118558": "you can check the data augmentation suggested by hoyt in the [diabetic retinopathy][1] competition. \r\nit might be a good starting point.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/diabetic-retinopathy-detection/forums/t/12842/starter-code-for-leaderboard-score-of-0-46",
    "118577": "Nathaniel Shimoni: Thanks for the link.",
    "118582": "[quote=msram;118556]\r\n\r\nHave you used data augmentation before? If so, how well did it do? May be I will also try that.\r\n\r\n[/quote]\r\n\r\nYes, it did very well for me in the Right Whale Detection competition. In that competition I was using the FANN library (in Ruby), which is not competitive with CNNs. But with data augmentation plus ensembling the end result was pretty good.\r\n\r\nLike feature engineering, there is an element of domain knowledge that helps you guess what kind of changes are helpful. I'd guess a variety of small affine transformations are going to be helpful, and also  brightness/contrast.\r\n\r\nThe alternative to data augmentation, would be normalisation and/or automated cropping - if you can reliably extract the driver's head into a fixed pose form each image, then you have much less need to generate variations of the existing poses.",
    "118667": "Neil Slater: I see. I will try these things one by one. Thanks for the detailed response."
  },
  "source": "meta"
}