{
  "id": 20370,
  "title": "Model behavior on validation set error",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20370",
  "author_name": "",
  "post_date": "2016-04-23T20:07:49.347Z",
  "votes": null,
  "comment_count": 4,
  "views": 744,
  "content": "<p>When I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?</p>\n\n<hr>\n\n<h3>Edit</h3>\n\n<p>This is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label</p>\n\n<p>[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2</p>\n\n<p>[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\n  0.13907903  0.07815139  0.11678198  0.13071878] c9</p>\n\n<p>[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\n  0.13587254  0.08120919  0.11257692  0.13526757] c5</p>\n\n<p>[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\n  0.13984248  0.07813746  0.1157885   0.13624895] c0</p>\n\n<p>[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\n  0.13927543  0.07820505  0.11453225  0.13505365] c0</p>",
  "messages": [
    {
      "id": "116380",
      "postDate": "04/23/2016 20:07:49",
      "content": "<p>When I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?</p>\n\n<hr>\n\n<h3>Edit</h3>\n\n<p>This is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label</p>\n\n<p>[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2</p>\n\n<p>[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\n  0.13907903  0.07815139  0.11678198  0.13071878] c9</p>\n\n<p>[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\n  0.13587254  0.08120919  0.11257692  0.13526757] c5</p>\n\n<p>[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\n  0.13984248  0.07813746  0.1157885   0.13624895] c0</p>\n\n<p>[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\n  0.13927543  0.07820505  0.11453225  0.13505365] c0</p>",
      "rawMarkdown": "When I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\r\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\r\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?\r\n________________________\r\n### Edit\r\nThis is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label\r\n\r\n[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\r\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2\r\n\r\n[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\r\n  0.13907903  0.07815139  0.11678198  0.13071878] c9\r\n\r\n[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\r\n  0.13587254  0.08120919  0.11257692  0.13526757] c5\r\n\r\n[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\r\n  0.13984248  0.07813746  0.1157885   0.13624895] c0\r\n\r\n[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\r\n  0.13927543  0.07820505  0.11453225  0.13505365] c0",
      "votes": null
    },
    {
      "id": "116418",
      "postDate": "04/24/2016 03:29:48",
      "content": "<p>[quote=Matteo Presutto;116380]</p>\n\n<p>When I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?</p>\n\n<hr>\n\n<h3>Edit</h3>\n\n<p>This is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label</p>\n\n<p>[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2</p>\n\n<p>[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\n  0.13907903  0.07815139  0.11678198  0.13071878] c9</p>\n\n<p>[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\n  0.13587254  0.08120919  0.11257692  0.13526757] c5</p>\n\n<p>[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\n  0.13984248  0.07813746  0.1157885   0.13624895] c0</p>\n\n<p>[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\n  0.13927543  0.07820505  0.11453225  0.13505365] c0</p>\n\n<p>[/quote]\nCheck out my post of this competition</p>",
      "rawMarkdown": "[quote=Matteo Presutto;116380]\r\n\r\nWhen I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\r\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\r\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?\r\n________________________\r\n### Edit\r\nThis is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label\r\n\r\n[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\r\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2\r\n\r\n[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\r\n  0.13907903  0.07815139  0.11678198  0.13071878] c9\r\n\r\n[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\r\n  0.13587254  0.08120919  0.11257692  0.13526757] c5\r\n\r\n[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\r\n  0.13984248  0.07813746  0.1157885   0.13624895] c0\r\n\r\n[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\r\n  0.13927543  0.07820505  0.11453225  0.13505365] c0\r\n\r\n[/quote]\r\nCheck out my post of this competition",
      "votes": null
    },
    {
      "id": "116437",
      "postDate": "04/24/2016 08:55:34",
      "content": "<p>I tried different learning rates, that seems not to be the problem. This is my architecture</p>\n\n<p>P.s: this is the output of a few test images after 6 epochs with learning rate 0.0005, in my previous post I was using 0.01 with momentum 0.9 </p>\n\n<p>[ 0.00669828  0.04828245  0.04955705  0.00644134  0.01015031  0.05190837\n  0.12816651  0.19230273  0.35769728  0.14879562] 9</p>\n\n<p>[ 0.02537988  0.06342643  0.05609277  0.01768731  0.019713    0.08014338\n  0.17849815  0.24909474  0.17819338  0.13177091] 5</p>\n\n<p>[ 0.01412927  0.04991864  0.05674135  0.00561293  0.01303112  0.1298552\n  0.14167444  0.1432393   0.16491506  0.28088278] 0</p>\n\n<p>[ 0.01534524  0.05045921  0.05817921  0.01287215  0.01433749  0.08339859\n  0.1626195   0.28111762  0.1971958   0.12447526] 0</p>\n\n<p>60it [00:26,  2.23it/s]</p>\n\n<p>Epoch 6 of 100 took 236.328s</p>\n\n<p>training loss:                2.323097</p>\n\n<p>validation loss:              2.525241</p>\n\n<p>validation accuracy:          11.83 %</p>",
      "rawMarkdown": "I tried different learning rates, that seems not to be the problem. This is my architecture\r\n\r\nP.s: this is the output of a few test images after 6 epochs with learning rate 0.0005, in my previous post I was using 0.01 with momentum 0.9 \r\n\r\n[ 0.00669828  0.04828245  0.04955705  0.00644134  0.01015031  0.05190837\r\n  0.12816651  0.19230273  0.35769728  0.14879562] 9\r\n\r\n[ 0.02537988  0.06342643  0.05609277  0.01768731  0.019713    0.08014338\r\n  0.17849815  0.24909474  0.17819338  0.13177091] 5\r\n\r\n[ 0.01412927  0.04991864  0.05674135  0.00561293  0.01303112  0.1298552\r\n  0.14167444  0.1432393   0.16491506  0.28088278] 0\r\n\r\n[ 0.01534524  0.05045921  0.05817921  0.01287215  0.01433749  0.08339859\r\n  0.1626195   0.28111762  0.1971958   0.12447526] 0\r\n\r\n60it [00:26,  2.23it/s]\r\n\r\nEpoch 6 of 100 took 236.328s\r\n  \r\ntraining loss:                2.323097\r\n  \r\nvalidation loss:              2.525241\r\n  \r\nvalidation accuracy:          11.83 %",
      "votes": null
    },
    {
      "id": "116453",
      "postDate": "04/24/2016 12:04:50",
      "content": "<p>Of course is not the learning rate's trouble, I am still overcoming it</p>",
      "rawMarkdown": "Of course is not the learning rate's trouble, I am still overcoming it",
      "votes": null
    },
    {
      "id": "116895",
      "postDate": "04/26/2016 13:53:01",
      "content": "<p>For completeness, I've found the problem. It was a very silly one (embarassing I would say), I forgot to shuffle the training set. In fact, as you can see, the first class gets very low prediction probability while the last ones are higher. The training error was not affected since the net was still able to overfit the set, I'm very surprised that a smaller network has been able to learn and generalize though!</p>",
      "rawMarkdown": "For completeness, I've found the problem. It was a very silly one (embarassing I would say), I forgot to shuffle the training set. In fact, as you can see, the first class gets very low prediction probability while the last ones are higher. The training error was not affected since the net was still able to overfit the set, I'm very surprised that a smaller network has been able to learn and generalize though!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 116418,
      "author_name": "meanku",
      "author_url": "",
      "post_date": "04/24/2016 03:29:48",
      "content": "<p>[quote=Matteo Presutto;116380]</p>\n\n<p>When I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?</p>\n\n<hr>\n\n<h3>Edit</h3>\n\n<p>This is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label</p>\n\n<p>[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2</p>\n\n<p>[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\n  0.13907903  0.07815139  0.11678198  0.13071878] c9</p>\n\n<p>[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\n  0.13587254  0.08120919  0.11257692  0.13526757] c5</p>\n\n<p>[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\n  0.13984248  0.07813746  0.1157885   0.13624895] c0</p>\n\n<p>[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\n  0.13927543  0.07820505  0.11453225  0.13505365] c0</p>\n\n<p>[/quote]\nCheck out my post of this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116437,
      "author_name": "matteopresutto",
      "author_url": "",
      "post_date": "04/24/2016 08:55:34",
      "content": "<p>I tried different learning rates, that seems not to be the problem. This is my architecture</p>\n\n<p>P.s: this is the output of a few test images after 6 epochs with learning rate 0.0005, in my previous post I was using 0.01 with momentum 0.9 </p>\n\n<p>[ 0.00669828  0.04828245  0.04955705  0.00644134  0.01015031  0.05190837\n  0.12816651  0.19230273  0.35769728  0.14879562] 9</p>\n\n<p>[ 0.02537988  0.06342643  0.05609277  0.01768731  0.019713    0.08014338\n  0.17849815  0.24909474  0.17819338  0.13177091] 5</p>\n\n<p>[ 0.01412927  0.04991864  0.05674135  0.00561293  0.01303112  0.1298552\n  0.14167444  0.1432393   0.16491506  0.28088278] 0</p>\n\n<p>[ 0.01534524  0.05045921  0.05817921  0.01287215  0.01433749  0.08339859\n  0.1626195   0.28111762  0.1971958   0.12447526] 0</p>\n\n<p>60it [00:26,  2.23it/s]</p>\n\n<p>Epoch 6 of 100 took 236.328s</p>\n\n<p>training loss:                2.323097</p>\n\n<p>validation loss:              2.525241</p>\n\n<p>validation accuracy:          11.83 %</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116453,
      "author_name": "meanku",
      "author_url": "",
      "post_date": "04/24/2016 12:04:50",
      "content": "<p>Of course is not the learning rate's trouble, I am still overcoming it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 116895,
      "author_name": "matteopresutto",
      "author_url": "",
      "post_date": "04/26/2016 13:53:01",
      "content": "<p>For completeness, I've found the problem. It was a very silly one (embarassing I would say), I forgot to shuffle the training set. In fact, as you can see, the first class gets very low prediction probability while the last ones are higher. The training error was not affected since the net was still able to overfit the set, I'm very surprised that a smaller network has been able to learn and generalize though!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116380": "When I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\r\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\r\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?\r\n________________________\r\n### Edit\r\nThis is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label\r\n\r\n[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\r\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2\r\n\r\n[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\r\n  0.13907903  0.07815139  0.11678198  0.13071878] c9\r\n\r\n[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\r\n  0.13587254  0.08120919  0.11257692  0.13526757] c5\r\n\r\n[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\r\n  0.13984248  0.07813746  0.1157885   0.13624895] c0\r\n\r\n[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\r\n  0.13927543  0.07820505  0.11453225  0.13505365] c0",
    "116418": "[quote=Matteo Presutto;116380]\r\n\r\nWhen I train deep CNN (more than 6 layers deep) a strange thing happens: the in-sample error decreases reaching eventually near zero values while the out-of-sample error doesn't improve at all.\r\nThis is how a model behaves when it overfits the training set but I'm using dropout after every layer, I even added the l2 regularization factor but nothing has changed.\r\nSince the network can overfit data I assume this is not a vanishing gradient problem, why does this happen? Did any of you experience the same thing?\r\n________________________\r\n### Edit\r\nThis is a sample of what the network outputs, the vector is the trained CNN output of an out-of-sample image (ordered from c0 to c9), at its right there's the true label\r\n\r\n[ 0.07908576  0.08674794  0.09045026  0.12157904  0.07243737  0.08600142\r\n  0.13547963  0.08158217  0.1121449   0.1344915 ]  c2\r\n\r\n[ 0.07588011  0.08523834  0.08786979  0.12765758  0.06849513  0.09012776\r\n  0.13907903  0.07815139  0.11678198  0.13071878] c9\r\n\r\n[ 0.08017623  0.08709945  0.09083226  0.12020027  0.0722551   0.08451048\r\n  0.13587254  0.08120919  0.11257692  0.13526757] c5\r\n\r\n[ 0.07580353  0.08478168  0.08867965  0.12542175  0.06753285  0.08776323\r\n  0.13984248  0.07813746  0.1157885   0.13624895] c0\r\n\r\n[ 0.07835174  0.08477517  0.08929344  0.12438807  0.07052168  0.08560355\r\n  0.13927543  0.07820505  0.11453225  0.13505365] c0\r\n\r\n[/quote]\r\nCheck out my post of this competition",
    "116437": "I tried different learning rates, that seems not to be the problem. This is my architecture\r\n\r\nP.s: this is the output of a few test images after 6 epochs with learning rate 0.0005, in my previous post I was using 0.01 with momentum 0.9 \r\n\r\n[ 0.00669828  0.04828245  0.04955705  0.00644134  0.01015031  0.05190837\r\n  0.12816651  0.19230273  0.35769728  0.14879562] 9\r\n\r\n[ 0.02537988  0.06342643  0.05609277  0.01768731  0.019713    0.08014338\r\n  0.17849815  0.24909474  0.17819338  0.13177091] 5\r\n\r\n[ 0.01412927  0.04991864  0.05674135  0.00561293  0.01303112  0.1298552\r\n  0.14167444  0.1432393   0.16491506  0.28088278] 0\r\n\r\n[ 0.01534524  0.05045921  0.05817921  0.01287215  0.01433749  0.08339859\r\n  0.1626195   0.28111762  0.1971958   0.12447526] 0\r\n\r\n60it [00:26,  2.23it/s]\r\n\r\nEpoch 6 of 100 took 236.328s\r\n  \r\ntraining loss:                2.323097\r\n  \r\nvalidation loss:              2.525241\r\n  \r\nvalidation accuracy:          11.83 %",
    "116453": "Of course is not the learning rate's trouble, I am still overcoming it",
    "116895": "For completeness, I've found the problem. It was a very silly one (embarassing I would say), I forgot to shuffle the training set. In fact, as you can see, the first class gets very low prediction probability while the last ones are higher. The training error was not affected since the net was still able to overfit the set, I'm very surprised that a smaller network has been able to learn and generalize though!"
  },
  "source": "meta"
}