{
  "id": 22614,
  "title": "How to do cross validation?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22614",
  "author_name": "",
  "post_date": "2016-08-02T00:09:50.940Z",
  "votes": 3,
  "comment_count": 8,
  "views": 782,
  "content": "<p>Congratulations to winners!</p>\n\n<p>Competition is over, and now it is time to share ideas. </p>\n\n<p>My biggest problem here was that I did not figure out how to do reliable cross-validation for this competition and as a result, I just overfitted Public Leaderboard.</p>\n\n<p>I am assuming that people that had the same result at Public and Private Leaderboards have some sacred knowledge about cross-validation techniques that allowed them to validate their scores locally.</p>\n\n<p>Can anyone share what they did?</p>",
  "messages": [
    {
      "id": "129723",
      "postDate": "08/02/2016 00:09:50",
      "content": "<p>Congratulations to winners!</p>\n\n<p>Competition is over, and now it is time to share ideas. </p>\n\n<p>My biggest problem here was that I did not figure out how to do reliable cross-validation for this competition and as a result, I just overfitted Public Leaderboard.</p>\n\n<p>I am assuming that people that had the same result at Public and Private Leaderboards have some sacred knowledge about cross-validation techniques that allowed them to validate their scores locally.</p>\n\n<p>Can anyone share what they did?</p>",
      "rawMarkdown": "Congratulations to winners!\r\n\r\nCompetition is over, and now it is time to share ideas. \r\n\r\nMy biggest problem here was that I did not figure out how to do reliable cross-validation for this competition and as a result, I just overfitted Public Leaderboard.\r\n\r\n\r\nI am assuming that people that had the same result at Public and Private Leaderboards have some sacred knowledge about cross-validation techniques that allowed them to validate their scores locally.\r\n\r\nCan anyone share what they did?",
      "votes": null
    },
    {
      "id": "129879",
      "postDate": "08/02/2016 17:45:46",
      "content": "<p>hi, here is my results:\npublic = 0.16820 (rank 37), private = 0.16522 (rank 20)</p>\n\n<p>cross-validation dataset: divide train image by {driver,action} clusters. Use 70 to 80% of clusters for training and the rest for validation.</p>\n\n<p>regularisation: this is very important. I try to make training difficult by synthesizing hard train samples.\nhard train sample = image + &quot;cut and paste&quot; of some region of another image of another class</p>\n\n<p>ensemble: 25 CNNs of googlenet_v1/v2, vgg_16/19, resnet50</p>",
      "rawMarkdown": "hi, here is my results:\r\npublic = 0.16820 (rank 37), private = 0.16522 (rank 20)\r\n\r\ncross-validation dataset: divide train image by {driver,action} clusters. Use 70 to 80% of clusters for training and the rest for validation.\r\n\r\nregularisation: this is very important. I try to make training difficult by synthesizing hard train samples.\r\nhard train sample = image + \"cut and paste\" of some region of another image of another class\r\n\r\nensemble: 25 CNNs of googlenet_v1/v2, vgg_16/19, resnet50",
      "votes": null
    },
    {
      "id": "129880",
      "postDate": "08/02/2016 17:49:06",
      "content": "<p>To do a reliable CV you just need to be sure that the same drivers you are using for training are not in the validation (holdout) fold.     </p>",
      "rawMarkdown": "To do a reliable CV you just need to be sure that the same drivers you are using for training are not in the validation (holdout) fold.",
      "votes": null
    },
    {
      "id": "129891",
      "postDate": "08/02/2016 18:42:14",
      "content": "<p>The inconsistency between CV and LB results have been bothering me too. I used 26-fold (based on drivers) CV, and only found out that I need to training longer than was suggested by the CV results to obtain a decent LB score. So I ended up using all the training data to build models and solely relied on the public LB for validation. To alleviate overfitting to the public LB, I trained 150+ models, as I believe the large ensemble is less likely to overfit. </p>\n\n<p>Public LB:  0.15592</p>\n\n<p>Private LB: 0.15985</p>",
      "rawMarkdown": "The inconsistency between CV and LB results have been bothering me too. I used 26-fold (based on drivers) CV, and only found out that I need to training longer than was suggested by the CV results to obtain a decent LB score. So I ended up using all the training data to build models and solely relied on the public LB for validation. To alleviate overfitting to the public LB, I trained 150+ models, as I believe the large ensemble is less likely to overfit. \r\n\r\nPublic LB:  0.15592\r\n\r\nPrivate LB: 0.15985",
      "votes": null
    },
    {
      "id": "129928",
      "postDate": "08/02/2016 22:12:45",
      "content": "<p>I started by &quot;leave one driver&quot; out (or 26 folds based on driver id) and cannot get reliable CV. (validation loss ranged from 0.001 to 0.7, some models stop at 6 epochs and some only have 1 epochs ). </p>\n\n<p>So I think I should increase the size of hold out ,like use 5 fold or even 3 folds based on driver id.</p>\n\n<p>Anyway I didn't have time to try that so I end up using 5 fold based on random split.  With ensemble  ,the public LB is 0.20979  and private LB:  0.20367</p>\n\n<p>I suppose for this dataset the optimum parameters of tuning a VGG19 is more or less the same no matter CV by driver_id or by random split.</p>",
      "rawMarkdown": "I started by \"leave one driver\" out (or 26 folds based on driver id) and cannot get reliable CV. (validation loss ranged from 0.001 to 0.7, some models stop at 6 epochs and some only have 1 epochs ). \r\n\r\nSo I think I should increase the size of hold out ,like use 5 fold or even 3 folds based on driver id.\r\n\r\nAnyway I didn't have time to try that so I end up using 5 fold based on random split.  With ensemble  ,the public LB is 0.20979  and private LB:  0.20367\r\n\r\n\r\nI suppose for this dataset the optimum parameters of tuning a VGG19 is more or less the same no matter CV by driver_id or by random split.",
      "votes": null
    },
    {
      "id": "129934",
      "postDate": "08/02/2016 23:10:55",
      "content": "<p>The right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. </p>\n\n<p>Here is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.</p>\n\n<p>Here is a <a href=\"https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\">link to the Kernel</a></p>\n\n<pre><code>## Simulating Small Data Set Bias and Variance\n# Jeff Hebert 8/2/2016\n\ncnt &lt;- 26\nset.seed(123)\nx &lt;- runif(cnt,0,10)\ny &lt;- x + rnorm(cnt, mean = 0, sd = 1)\nfolds &lt;- as.integer(cut(1:cnt, 4))\n\nplot(x,y, col = folds, pch = 19)\n\nmodel &lt;- lm(y ~ x)\nabline(model, lwd = 3)\n\nfor(fold in folds){\n  idx &lt;- folds == fold\n  x_tr &lt;- x[idx]\n  y_tr &lt;- y[idx]\n  x_te &lt;- x[-idx]\n  y_te &lt;- y[-idx]\n\n  model &lt;- lm(y_tr ~ x_tr)\n  abline(model, col = fold)\n}\n</code></pre>\n\n<p><img src=\"https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png\" alt=\"Simulating Small Data Set Bias and Variance\" title></p>",
      "rawMarkdown": "The right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. \r\n\r\nHere is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.\r\n\r\nHere is a [link to the Kernel][1]\r\n\r\n\r\n    ## Simulating Small Data Set Bias and Variance\r\n    # Jeff Hebert 8/2/2016\r\n    \r\n    cnt <- 26\r\n    set.seed(123)\r\n    x <- runif(cnt,0,10)\r\n    y <- x + rnorm(cnt, mean = 0, sd = 1)\r\n    folds <- as.integer(cut(1:cnt, 4))\r\n    \r\n    plot(x,y, col = folds, pch = 19)\r\n    \r\n    model <- lm(y ~ x)\r\n    abline(model, lwd = 3)\r\n    \r\n    for(fold in folds){\r\n      idx <- folds == fold\r\n      x_tr <- x[idx]\r\n      y_tr <- y[idx]\r\n      x_te <- x[-idx]\r\n      y_te <- y[-idx]\r\n      \r\n      model <- lm(y_tr ~ x_tr)\r\n      abline(model, col = fold)\r\n    }\r\n\r\n![Simulating Small Data Set Bias and Variance][2]\r\n\r\n\r\n  [1]: https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\r\n  [2]: https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png",
      "votes": null
    },
    {
      "id": "129936",
      "postDate": "08/02/2016 23:46:03",
      "content": "<p>My takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.</p>",
      "rawMarkdown": "My takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.",
      "votes": null
    },
    {
      "id": "129940",
      "postDate": "08/03/2016 00:25:24",
      "content": "<p>[quote=kyv;129936]</p>\n\n<p>My takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.</p>\n\n<p>[/quote]\nI had hope for the same assumption(Basically we mimic how Random Forest works) and it affected me in a way:</p>\n\n<p>0.18 Public =&gt; 0.2 Private</p>",
      "rawMarkdown": "[quote=kyv;129936]\r\n\r\nMy takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.\r\n\r\n[/quote]\r\nI had hope for the same assumption(Basically we mimic how Random Forest works) and it affected me in a way:\r\n\r\n0.18 Public => 0.2 Private",
      "votes": null
    },
    {
      "id": "129942",
      "postDate": "08/03/2016 00:35:12",
      "content": "<p>[quote=JeffH;129934]</p>\n\n<p>The right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. </p>\n\n<p>Here is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.</p>\n\n<p>Here is a <a href=\"https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\">link to the Kernel</a></p>\n\n<pre><code>## Simulating Small Data Set Bias and Variance\n# Jeff Hebert 8/2/2016\n\ncnt &lt;- 26\nset.seed(123)\nx &lt;- runif(cnt,0,10)\ny &lt;- x + rnorm(cnt, mean = 0, sd = 1)\nfolds &lt;- as.integer(cut(1:cnt, 4))\n\nplot(x,y, col = folds, pch = 19)\n\nmodel &lt;- lm(y ~ x)\nabline(model, lwd = 3)\n\nfor(fold in folds){\n  idx &lt;- folds == fold\n  x_tr &lt;- x[idx]\n  y_tr &lt;- y[idx]\n  x_te &lt;- x[-idx]\n  y_te &lt;- y[-idx]\n\n  model &lt;- lm(y_tr ~ x_tr)\n  abline(model, col = fold)\n}\n</code></pre>\n\n<p><img src=\"https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png\" alt=\"Simulating Small Data Set Bias and Variance\" title></p>\n\n<p>[/quote]</p>\n\n<p>I have some different opinion.</p>\n\n<ul>\n<li>CV is for model selection, ensemble reduces variance.</li>\n<li>Re-training on the whole dataset reduces variance for a single model, but the effectiveness of model averaging could suffer due to more correlated predictions.</li>\n</ul>",
      "rawMarkdown": "[quote=JeffH;129934]\r\n\r\nThe right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. \r\n\r\nHere is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.\r\n\r\nHere is a [link to the Kernel][1]\r\n\r\n\r\n    ## Simulating Small Data Set Bias and Variance\r\n    # Jeff Hebert 8/2/2016\r\n    \r\n    cnt <- 26\r\n    set.seed(123)\r\n    x <- runif(cnt,0,10)\r\n    y <- x + rnorm(cnt, mean = 0, sd = 1)\r\n    folds <- as.integer(cut(1:cnt, 4))\r\n    \r\n    plot(x,y, col = folds, pch = 19)\r\n    \r\n    model <- lm(y ~ x)\r\n    abline(model, lwd = 3)\r\n    \r\n    for(fold in folds){\r\n      idx <- folds == fold\r\n      x_tr <- x[idx]\r\n      y_tr <- y[idx]\r\n      x_te <- x[-idx]\r\n      y_te <- y[-idx]\r\n      \r\n      model <- lm(y_tr ~ x_tr)\r\n      abline(model, col = fold)\r\n    }\r\n\r\n![Simulating Small Data Set Bias and Variance][2]\r\n\r\n\r\n  [1]: https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\r\n  [2]: https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png\r\n\r\n[/quote]\r\n\r\nI have some different opinion.\r\n\r\n - CV is for model selection, ensemble reduces variance.\r\n - Re-training on the whole dataset reduces variance for a single model, but the effectiveness of model averaging could suffer due to more correlated predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129879,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/02/2016 17:45:46",
      "content": "<p>hi, here is my results:\npublic = 0.16820 (rank 37), private = 0.16522 (rank 20)</p>\n\n<p>cross-validation dataset: divide train image by {driver,action} clusters. Use 70 to 80% of clusters for training and the rest for validation.</p>\n\n<p>regularisation: this is very important. I try to make training difficult by synthesizing hard train samples.\nhard train sample = image + &quot;cut and paste&quot; of some region of another image of another class</p>\n\n<p>ensemble: 25 CNNs of googlenet_v1/v2, vgg_16/19, resnet50</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129880,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "08/02/2016 17:49:06",
      "content": "<p>To do a reliable CV you just need to be sure that the same drivers you are using for training are not in the validation (holdout) fold.     </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129891,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "08/02/2016 18:42:14",
      "content": "<p>The inconsistency between CV and LB results have been bothering me too. I used 26-fold (based on drivers) CV, and only found out that I need to training longer than was suggested by the CV results to obtain a decent LB score. So I ended up using all the training data to build models and solely relied on the public LB for validation. To alleviate overfitting to the public LB, I trained 150+ models, as I believe the large ensemble is less likely to overfit. </p>\n\n<p>Public LB:  0.15592</p>\n\n<p>Private LB: 0.15985</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129928,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "08/02/2016 22:12:45",
      "content": "<p>I started by &quot;leave one driver&quot; out (or 26 folds based on driver id) and cannot get reliable CV. (validation loss ranged from 0.001 to 0.7, some models stop at 6 epochs and some only have 1 epochs ). </p>\n\n<p>So I think I should increase the size of hold out ,like use 5 fold or even 3 folds based on driver id.</p>\n\n<p>Anyway I didn't have time to try that so I end up using 5 fold based on random split.  With ensemble  ,the public LB is 0.20979  and private LB:  0.20367</p>\n\n<p>I suppose for this dataset the optimum parameters of tuning a VGG19 is more or less the same no matter CV by driver_id or by random split.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129934,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "08/02/2016 23:10:55",
      "content": "<p>The right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. </p>\n\n<p>Here is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.</p>\n\n<p>Here is a <a href=\"https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\">link to the Kernel</a></p>\n\n<pre><code>## Simulating Small Data Set Bias and Variance\n# Jeff Hebert 8/2/2016\n\ncnt &lt;- 26\nset.seed(123)\nx &lt;- runif(cnt,0,10)\ny &lt;- x + rnorm(cnt, mean = 0, sd = 1)\nfolds &lt;- as.integer(cut(1:cnt, 4))\n\nplot(x,y, col = folds, pch = 19)\n\nmodel &lt;- lm(y ~ x)\nabline(model, lwd = 3)\n\nfor(fold in folds){\n  idx &lt;- folds == fold\n  x_tr &lt;- x[idx]\n  y_tr &lt;- y[idx]\n  x_te &lt;- x[-idx]\n  y_te &lt;- y[-idx]\n\n  model &lt;- lm(y_tr ~ x_tr)\n  abline(model, col = fold)\n}\n</code></pre>\n\n<p><img src=\"https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png\" alt=\"Simulating Small Data Set Bias and Variance\" title></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129936,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/02/2016 23:46:03",
      "content": "<p>My takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129940,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "08/03/2016 00:25:24",
      "content": "<p>[quote=kyv;129936]</p>\n\n<p>My takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.</p>\n\n<p>[/quote]\nI had hope for the same assumption(Basically we mimic how Random Forest works) and it affected me in a way:</p>\n\n<p>0.18 Public =&gt; 0.2 Private</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129942,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "08/03/2016 00:35:12",
      "content": "<p>[quote=JeffH;129934]</p>\n\n<p>The right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. </p>\n\n<p>Here is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.</p>\n\n<p>Here is a <a href=\"https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\">link to the Kernel</a></p>\n\n<pre><code>## Simulating Small Data Set Bias and Variance\n# Jeff Hebert 8/2/2016\n\ncnt &lt;- 26\nset.seed(123)\nx &lt;- runif(cnt,0,10)\ny &lt;- x + rnorm(cnt, mean = 0, sd = 1)\nfolds &lt;- as.integer(cut(1:cnt, 4))\n\nplot(x,y, col = folds, pch = 19)\n\nmodel &lt;- lm(y ~ x)\nabline(model, lwd = 3)\n\nfor(fold in folds){\n  idx &lt;- folds == fold\n  x_tr &lt;- x[idx]\n  y_tr &lt;- y[idx]\n  x_te &lt;- x[-idx]\n  y_te &lt;- y[-idx]\n\n  model &lt;- lm(y_tr ~ x_tr)\n  abline(model, col = fold)\n}\n</code></pre>\n\n<p><img src=\"https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png\" alt=\"Simulating Small Data Set Bias and Variance\" title></p>\n\n<p>[/quote]</p>\n\n<p>I have some different opinion.</p>\n\n<ul>\n<li>CV is for model selection, ensemble reduces variance.</li>\n<li>Re-training on the whole dataset reduces variance for a single model, but the effectiveness of model averaging could suffer due to more correlated predictions.</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129723": "Congratulations to winners!\r\n\r\nCompetition is over, and now it is time to share ideas. \r\n\r\nMy biggest problem here was that I did not figure out how to do reliable cross-validation for this competition and as a result, I just overfitted Public Leaderboard.\r\n\r\n\r\nI am assuming that people that had the same result at Public and Private Leaderboards have some sacred knowledge about cross-validation techniques that allowed them to validate their scores locally.\r\n\r\nCan anyone share what they did?",
    "129879": "hi, here is my results:\r\npublic = 0.16820 (rank 37), private = 0.16522 (rank 20)\r\n\r\ncross-validation dataset: divide train image by {driver,action} clusters. Use 70 to 80% of clusters for training and the rest for validation.\r\n\r\nregularisation: this is very important. I try to make training difficult by synthesizing hard train samples.\r\nhard train sample = image + \"cut and paste\" of some region of another image of another class\r\n\r\nensemble: 25 CNNs of googlenet_v1/v2, vgg_16/19, resnet50",
    "129880": "To do a reliable CV you just need to be sure that the same drivers you are using for training are not in the validation (holdout) fold.",
    "129891": "The inconsistency between CV and LB results have been bothering me too. I used 26-fold (based on drivers) CV, and only found out that I need to training longer than was suggested by the CV results to obtain a decent LB score. So I ended up using all the training data to build models and solely relied on the public LB for validation. To alleviate overfitting to the public LB, I trained 150+ models, as I believe the large ensemble is less likely to overfit. \r\n\r\nPublic LB:  0.15592\r\n\r\nPrivate LB: 0.15985",
    "129928": "I started by \"leave one driver\" out (or 26 folds based on driver id) and cannot get reliable CV. (validation loss ranged from 0.001 to 0.7, some models stop at 6 epochs and some only have 1 epochs ). \r\n\r\nSo I think I should increase the size of hold out ,like use 5 fold or even 3 folds based on driver id.\r\n\r\nAnyway I didn't have time to try that so I end up using 5 fold based on random split.  With ensemble  ,the public LB is 0.20979  and private LB:  0.20367\r\n\r\n\r\nI suppose for this dataset the optimum parameters of tuning a VGG19 is more or less the same no matter CV by driver_id or by random split.",
    "129934": "The right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. \r\n\r\nHere is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.\r\n\r\nHere is a [link to the Kernel][1]\r\n\r\n\r\n    ## Simulating Small Data Set Bias and Variance\r\n    # Jeff Hebert 8/2/2016\r\n    \r\n    cnt <- 26\r\n    set.seed(123)\r\n    x <- runif(cnt,0,10)\r\n    y <- x + rnorm(cnt, mean = 0, sd = 1)\r\n    folds <- as.integer(cut(1:cnt, 4))\r\n    \r\n    plot(x,y, col = folds, pch = 19)\r\n    \r\n    model <- lm(y ~ x)\r\n    abline(model, lwd = 3)\r\n    \r\n    for(fold in folds){\r\n      idx <- folds == fold\r\n      x_tr <- x[idx]\r\n      y_tr <- y[idx]\r\n      x_te <- x[-idx]\r\n      y_te <- y[-idx]\r\n      \r\n      model <- lm(y_tr ~ x_tr)\r\n      abline(model, col = fold)\r\n    }\r\n\r\n![Simulating Small Data Set Bias and Variance][2]\r\n\r\n\r\n  [1]: https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\r\n  [2]: https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png",
    "129936": "My takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.",
    "129940": "[quote=kyv;129936]\r\n\r\nMy takeaway was that overffiting is OK as long as models overfitted on different examples and then averaged.  I used 8-fold CV by drivers. Difference between 0.33 on single fold model and 0.20 must mean overfitting was completely removed by avereging.\r\n\r\n[/quote]\r\nI had hope for the same assumption(Basically we mimic how Random Forest works) and it affected me in a way:\r\n\r\n0.18 Public => 0.2 Private",
    "129942": "[quote=JeffH;129934]\r\n\r\nThe right way to do CV for this competition is to split by driver. However, remember that CV is for model selection to reduce variance (so you get the same results with any new data). After you get a stable model, you should train it with the entire data set. Many people combine these two steps. This saves computing time, but can lead to increased bias (lower accuracy) because you are using subsets of the data to train. This competition was particularly difficult because we really only had   26 unique drivers. It seems that Heng CherKeng got around this by simulating new cases with cut-and-paste. Many people ended up randomly selecting from the images and tuning based on the public LB. This worked to some degree, but you definitely have the risk of over-fitting to the LB. \r\n\r\nHere is an R script and plot that demonstrates the effect of small sample size. Feel free to run yourself with different seeds.\r\n\r\nHere is a [link to the Kernel][1]\r\n\r\n\r\n    ## Simulating Small Data Set Bias and Variance\r\n    # Jeff Hebert 8/2/2016\r\n    \r\n    cnt <- 26\r\n    set.seed(123)\r\n    x <- runif(cnt,0,10)\r\n    y <- x + rnorm(cnt, mean = 0, sd = 1)\r\n    folds <- as.integer(cut(1:cnt, 4))\r\n    \r\n    plot(x,y, col = folds, pch = 19)\r\n    \r\n    model <- lm(y ~ x)\r\n    abline(model, lwd = 3)\r\n    \r\n    for(fold in folds){\r\n      idx <- folds == fold\r\n      x_tr <- x[idx]\r\n      y_tr <- y[idx]\r\n      x_te <- x[-idx]\r\n      y_te <- y[-idx]\r\n      \r\n      model <- lm(y_tr ~ x_tr)\r\n      abline(model, col = fold)\r\n    }\r\n\r\n![Simulating Small Data Set Bias and Variance][2]\r\n\r\n\r\n  [1]: https://www.kaggle.com/jeffhebert/state-farm-distracted-driver-detection/simulating-small-data-set\r\n  [2]: https://www.kaggle.io/svf/319705/219bd839a1c0f1f479e0ae4a4fadbf9e/Rplot001.png\r\n\r\n[/quote]\r\n\r\nI have some different opinion.\r\n\r\n - CV is for model selection, ensemble reduces variance.\r\n - Re-training on the whole dataset reduces variance for a single model, but the effectiveness of model averaging could suffer due to more correlated predictions."
  },
  "source": "meta"
}