{
  "id": 21117,
  "title": "Torch starter code",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/21117",
  "author_name": "",
  "post_date": "2016-05-21T04:48:41.287Z",
  "votes": 1,
  "comment_count": 2,
  "views": 978,
  "content": "<p>I'd like to share the Torch implementation that I've been working on. It's somewhat comparable to the Keras code that came out a while ago in that the main model is VGG  and trained from scratch. I thought I'd help along people who wanted to use Torch for any reason.</p>\n\n<p>For aggregating the models, I simply averaged the outputs. Doing the geometric mean didn't work for me because I get lots of values close to zero. If anyone has an explanation on why it would work I'd be interested in hearing it.</p>\n\n<p>This particular model was trained on 13 folds for 15 epoch (though 10 would have been enough). I used a GTX 960 and training took about an hour, with 13s/fold. The training loss was 1.72, validation loss 1.89 (though it just averaged the loss of the models, rather than ensembling them), LB score 1.55  (with model ensembling)</p>\n\n<p>This is the command I used to generate the submission:</p>\n\n<pre><code>th -i main.lua --batchSize 16 -r 0.003 --model vgg_small -h24 -w32 -d ps.t7 --max_epoch 15 -f 13 -t adam -r_factor 0.1 --lr_schedule 5 -m 0.1 -S -L2 0.001 -g\n</code></pre>\n\n<p>The parameters, in order, express batch size (obviously), learning rate, model (defined in models/vgg_small.lua), height, width, data file to use (if it's already created, don't use the -g flag), number of epochs (obviously), number of folds, training algorithm, learning rate decay factor, learning rate decay schedule (after this many epochs, multiply by lr_factor), momentum, generate submission (-S), L2 penalty, generate data (will be saved as ps.t7, omit -g if already saved). The submission file is generated as submission.csv.</p>\n\n<p>If you want a better score, you can change the resolution from 32x24 to 64x48, use a bigger model and 26 folds. Some model definitions are in the models/ folder. Note that the *_small.lua files are used for 32x24 images while the other ones are for 64x48.</p>\n\n<p>There are a number of things that can be improved, and I would welcome pull requests, though I am going to continue working on this.</p>\n\n<p>You can find the repo here: <a href=\"https://github.com/lucianionita/distracted-drivers-torch\">https://github.com/lucianionita/distracted-drivers-torch</a>                                                             </p>",
  "messages": [
    {
      "id": "120869",
      "postDate": "05/21/2016 04:48:41",
      "content": "<p>I'd like to share the Torch implementation that I've been working on. It's somewhat comparable to the Keras code that came out a while ago in that the main model is VGG  and trained from scratch. I thought I'd help along people who wanted to use Torch for any reason.</p>\n\n<p>For aggregating the models, I simply averaged the outputs. Doing the geometric mean didn't work for me because I get lots of values close to zero. If anyone has an explanation on why it would work I'd be interested in hearing it.</p>\n\n<p>This particular model was trained on 13 folds for 15 epoch (though 10 would have been enough). I used a GTX 960 and training took about an hour, with 13s/fold. The training loss was 1.72, validation loss 1.89 (though it just averaged the loss of the models, rather than ensembling them), LB score 1.55  (with model ensembling)</p>\n\n<p>This is the command I used to generate the submission:</p>\n\n<pre><code>th -i main.lua --batchSize 16 -r 0.003 --model vgg_small -h24 -w32 -d ps.t7 --max_epoch 15 -f 13 -t adam -r_factor 0.1 --lr_schedule 5 -m 0.1 -S -L2 0.001 -g\n</code></pre>\n\n<p>The parameters, in order, express batch size (obviously), learning rate, model (defined in models/vgg_small.lua), height, width, data file to use (if it's already created, don't use the -g flag), number of epochs (obviously), number of folds, training algorithm, learning rate decay factor, learning rate decay schedule (after this many epochs, multiply by lr_factor), momentum, generate submission (-S), L2 penalty, generate data (will be saved as ps.t7, omit -g if already saved). The submission file is generated as submission.csv.</p>\n\n<p>If you want a better score, you can change the resolution from 32x24 to 64x48, use a bigger model and 26 folds. Some model definitions are in the models/ folder. Note that the *_small.lua files are used for 32x24 images while the other ones are for 64x48.</p>\n\n<p>There are a number of things that can be improved, and I would welcome pull requests, though I am going to continue working on this.</p>\n\n<p>You can find the repo here: <a href=\"https://github.com/lucianionita/distracted-drivers-torch\">https://github.com/lucianionita/distracted-drivers-torch</a>                                                             </p>",
      "rawMarkdown": "I'd like to share the Torch implementation that I've been working on. It's somewhat comparable to the Keras code that came out a while ago in that the main model is VGG  and trained from scratch. I thought I'd help along people who wanted to use Torch for any reason.\r\n\r\nFor aggregating the models, I simply averaged the outputs. Doing the geometric mean didn't work for me because I get lots of values close to zero. If anyone has an explanation on why it would work I'd be interested in hearing it.\r\n\r\nThis particular model was trained on 13 folds for 15 epoch (though 10 would have been enough). I used a GTX 960 and training took about an hour, with 13s/fold. The training loss was 1.72, validation loss 1.89 (though it just averaged the loss of the models, rather than ensembling them), LB score 1.55  (with model ensembling)\r\n\r\n\r\n\r\nThis is the command I used to generate the submission:\r\n\r\n    th -i main.lua --batchSize 16 -r 0.003 --model vgg_small -h24 -w32 -d ps.t7 --max_epoch 15 -f 13 -t adam -r_factor 0.1 --lr_schedule 5 -m 0.1 -S -L2 0.001 -g\r\n\r\nThe parameters, in order, express batch size (obviously), learning rate, model (defined in models/vgg_small.lua), height, width, data file to use (if it's already created, don't use the -g flag), number of epochs (obviously), number of folds, training algorithm, learning rate decay factor, learning rate decay schedule (after this many epochs, multiply by lr_factor), momentum, generate submission (-S), L2 penalty, generate data (will be saved as ps.t7, omit -g if already saved). The submission file is generated as submission.csv.\r\n\r\nIf you want a better score, you can change the resolution from 32x24 to 64x48, use a bigger model and 26 folds. Some model definitions are in the models/ folder. Note that the *_small.lua files are used for 32x24 images while the other ones are for 64x48.\r\n\r\nThere are a number of things that can be improved, and I would welcome pull requests, though I am going to continue working on this.\r\n\r\n You can find the repo here: https://github.com/lucianionita/distracted-drivers-torch",
      "votes": null
    },
    {
      "id": "120891",
      "postDate": "05/21/2016 12:39:21",
      "content": "<p>hi, I notice that you use N-fold to train N models(and also in another VGG net), and take the average, can you explain the reason for this ? Is it because the big variance of NN models?</p>",
      "rawMarkdown": "hi, I notice that you use N-fold to train N models(and also in another VGG net), and take the average, can you explain the reason for this ? Is it because the big variance of NN models?",
      "votes": null
    },
    {
      "id": "121008",
      "postDate": "05/22/2016 18:10:23",
      "content": "<p>My main reason for doing this is that I want to get some feedback as the model is training. Most importantly I want to get a sense of how the training and validation losses change over time and in relation to each other. I haven't tried training a full model (with no hold-out set) yet, but I suspect that the variance induced by partitioning the training set during N-fold training does push the loss down a bit more than training one model on the entire dataset. </p>",
      "rawMarkdown": "My main reason for doing this is that I want to get some feedback as the model is training. Most importantly I want to get a sense of how the training and validation losses change over time and in relation to each other. I haven't tried training a full model (with no hold-out set) yet, but I suspect that the variance induced by partitioning the training set during N-fold training does push the loss down a bit more than training one model on the entire dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120891,
      "author_name": "zeliek",
      "author_url": "",
      "post_date": "05/21/2016 12:39:21",
      "content": "<p>hi, I notice that you use N-fold to train N models(and also in another VGG net), and take the average, can you explain the reason for this ? Is it because the big variance of NN models?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121008,
      "author_name": "ilucian",
      "author_url": "",
      "post_date": "05/22/2016 18:10:23",
      "content": "<p>My main reason for doing this is that I want to get some feedback as the model is training. Most importantly I want to get a sense of how the training and validation losses change over time and in relation to each other. I haven't tried training a full model (with no hold-out set) yet, but I suspect that the variance induced by partitioning the training set during N-fold training does push the loss down a bit more than training one model on the entire dataset. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "120869": "I'd like to share the Torch implementation that I've been working on. It's somewhat comparable to the Keras code that came out a while ago in that the main model is VGG  and trained from scratch. I thought I'd help along people who wanted to use Torch for any reason.\r\n\r\nFor aggregating the models, I simply averaged the outputs. Doing the geometric mean didn't work for me because I get lots of values close to zero. If anyone has an explanation on why it would work I'd be interested in hearing it.\r\n\r\nThis particular model was trained on 13 folds for 15 epoch (though 10 would have been enough). I used a GTX 960 and training took about an hour, with 13s/fold. The training loss was 1.72, validation loss 1.89 (though it just averaged the loss of the models, rather than ensembling them), LB score 1.55  (with model ensembling)\r\n\r\n\r\n\r\nThis is the command I used to generate the submission:\r\n\r\n    th -i main.lua --batchSize 16 -r 0.003 --model vgg_small -h24 -w32 -d ps.t7 --max_epoch 15 -f 13 -t adam -r_factor 0.1 --lr_schedule 5 -m 0.1 -S -L2 0.001 -g\r\n\r\nThe parameters, in order, express batch size (obviously), learning rate, model (defined in models/vgg_small.lua), height, width, data file to use (if it's already created, don't use the -g flag), number of epochs (obviously), number of folds, training algorithm, learning rate decay factor, learning rate decay schedule (after this many epochs, multiply by lr_factor), momentum, generate submission (-S), L2 penalty, generate data (will be saved as ps.t7, omit -g if already saved). The submission file is generated as submission.csv.\r\n\r\nIf you want a better score, you can change the resolution from 32x24 to 64x48, use a bigger model and 26 folds. Some model definitions are in the models/ folder. Note that the *_small.lua files are used for 32x24 images while the other ones are for 64x48.\r\n\r\nThere are a number of things that can be improved, and I would welcome pull requests, though I am going to continue working on this.\r\n\r\n You can find the repo here: https://github.com/lucianionita/distracted-drivers-torch",
    "120891": "hi, I notice that you use N-fold to train N models(and also in another VGG net), and take the average, can you explain the reason for this ? Is it because the big variance of NN models?",
    "121008": "My main reason for doing this is that I want to get some feedback as the model is training. Most importantly I want to get a sense of how the training and validation losses change over time and in relation to each other. I haven't tried training a full model (with no hold-out set) yet, but I suspect that the variance induced by partitioning the training set during N-fold training does push the loss down a bit more than training one model on the entire dataset."
  },
  "source": "meta"
}