{
  "id": 20561,
  "title": "Is it normal if testing loss is less than training loss?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20561",
  "author_name": "",
  "post_date": "2016-04-30T04:21:34.703Z",
  "votes": null,
  "comment_count": 6,
  "views": 1011,
  "content": "<p>I sometimes encountered this when trying different parameters. Is this normal? What are the possible reasons leading to this? Anyone has any idea?</p>",
  "messages": [
    {
      "id": "117665",
      "postDate": "04/30/2016 04:21:34",
      "content": "<p>I sometimes encountered this when trying different parameters. Is this normal? What are the possible reasons leading to this? Anyone has any idea?</p>",
      "rawMarkdown": "I sometimes encountered this when trying different parameters. Is this normal? What are the possible reasons leading to this? Anyone has any idea?",
      "votes": null
    },
    {
      "id": "117666",
      "postDate": "04/30/2016 04:32:06",
      "content": "<p>You might get lucky to find a set of parameters that happens to fit the distribution of the test data really well. I've seen it when using aggressive regularization. Does it happen consistently?</p>",
      "rawMarkdown": "You might get lucky to find a set of parameters that happens to fit the distribution of the test data really well. I've seen it when using aggressive regularization. Does it happen consistently?",
      "votes": null
    },
    {
      "id": "117668",
      "postDate": "04/30/2016 04:40:02",
      "content": "<p>@Lucian, actually not. For specific parameters, it is. But the thing is, the score of lb is not as expected. I don't know what happened. </p>",
      "rawMarkdown": "Lucian, actually not. For specific parameters, it is. But the thing is, the score of lb is not as expected. I don't know what happened.",
      "votes": null
    },
    {
      "id": "117669",
      "postDate": "04/30/2016 04:50:42",
      "content": "<p>Not as expected as in too high or too low?</p>",
      "rawMarkdown": "Not as expected as in too high or too low?",
      "votes": null
    },
    {
      "id": "117676",
      "postDate": "04/30/2016 06:21:24",
      "content": "<p>When splitting cv by driver, I have seen lower estimated loss for certain drivers consistently. Drivers<code>p045</code> and <code>p051</code> will consistently score lower than training loss, and it seems quite hard to overfit them just by using dropout for regularisation.</p>\n\n<p>One thing to check if it varies - are you always splitting your CV set the exact same way - i.e. same data in it each time (hopefully it is also split by driver)? Taking a different CV set each time could be creating this effect artificially for you since each driver can score very differently.</p>\n\n<p>I tracked this effect in some of my training. Here, the training loss is around 0.7 each time (not recorded, sorry), but the CV loss is known for each driver because I did 26-fold CV:</p>\n\n<pre><code>'p026' =&gt; 1.27, 1.65, 1.53\n'p047' =&gt; 1.33, 1.17, 1.07\n'p045' =&gt; 0.58, 0.51, 0.66\n'p022' =&gt; 1.42, 1.67, 1.31\n'p035' =&gt; 0.83, 0.95, 1.08\n'p024' =&gt; 1.21, 1.35, 1.45\n'p015' =&gt; 2.13, 2.59, 2.51\n'p061' =&gt; 1.21, 1.23, 1.27\n'p056' =&gt; 1.10, 1.03, 1.12\n'p021' =&gt; 1.02, 1.09, 1.19\n'p081' =&gt; 2.07, 1.88, 1.95\n'p039' =&gt; 1.51, 2.26, 1.56\n'p014' =&gt; 1.71, 1.78, 1.46\n'p075' =&gt; 2.00, 2.03, 1.96\n'p002' =&gt; 1.57, 1.72, 1.94\n'p049' =&gt; 1.12, 1.18, 1.24\n'p041' =&gt; 0.91, 0.86, 0.68\n'p016' =&gt; 1.21, 1.24, 1.42\n'p012' =&gt; 1.24, 1.41, 1.41\n'p066' =&gt; 1.70, 1.82, 1.69\n'p042' =&gt; 0.61, 0.63, 1.13\n'p051' =&gt; 0.53, 0.49, 0.56\n'p064' =&gt; 1.68, 1.52, 1.48\n'p052' =&gt; 0.88, 0.85, 0.71\n'p072' =&gt; 2.63, 6.35, 2.37\n'p050' =&gt; 2.27, 1.88, 1.93\n</code></pre>\n\n<p>The columns correspond to slight changes in architecture, and score 1.13, 1.14 and 1.17 respectively on the public leaderboard.</p>",
      "rawMarkdown": "When splitting cv by driver, I have seen lower estimated loss for certain drivers consistently. Drivers`p045` and `p051` will consistently score lower than training loss, and it seems quite hard to overfit them just by using dropout for regularisation.\r\n\r\nOne thing to check if it varies - are you always splitting your CV set the exact same way - i.e. same data in it each time (hopefully it is also split by driver)? Taking a different CV set each time could be creating this effect artificially for you since each driver can score very differently.\r\n\r\nI tracked this effect in some of my training. Here, the training loss is around 0.7 each time (not recorded, sorry), but the CV loss is known for each driver because I did 26-fold CV:\r\n    \r\n    'p026' => 1.27, 1.65, 1.53\r\n    'p047' => 1.33, 1.17, 1.07\r\n    'p045' => 0.58, 0.51, 0.66\r\n    'p022' => 1.42, 1.67, 1.31\r\n    'p035' => 0.83, 0.95, 1.08\r\n    'p024' => 1.21, 1.35, 1.45\r\n    'p015' => 2.13, 2.59, 2.51\r\n    'p061' => 1.21, 1.23, 1.27\r\n    'p056' => 1.10, 1.03, 1.12\r\n    'p021' => 1.02, 1.09, 1.19\r\n    'p081' => 2.07, 1.88, 1.95\r\n    'p039' => 1.51, 2.26, 1.56\r\n    'p014' => 1.71, 1.78, 1.46\r\n    'p075' => 2.00, 2.03, 1.96\r\n    'p002' => 1.57, 1.72, 1.94\r\n    'p049' => 1.12, 1.18, 1.24\r\n    'p041' => 0.91, 0.86, 0.68\r\n    'p016' => 1.21, 1.24, 1.42\r\n    'p012' => 1.24, 1.41, 1.41\r\n    'p066' => 1.70, 1.82, 1.69\r\n    'p042' => 0.61, 0.63, 1.13\r\n    'p051' => 0.53, 0.49, 0.56\r\n    'p064' => 1.68, 1.52, 1.48\r\n    'p052' => 0.88, 0.85, 0.71\r\n    'p072' => 2.63, 6.35, 2.37\r\n    'p050' => 2.27, 1.88, 1.93\r\n\r\nThe columns correspond to slight changes in architecture, and score 1.13, 1.14 and 1.17 respectively on the public leaderboard.",
      "votes": null
    },
    {
      "id": "117712",
      "postDate": "04/30/2016 13:50:25",
      "content": "<p>@Lucian, the lb is higher than the local cv.</p>\n\n<p>@Neil, I saw something like this too. Usually, I do 5-10 folder cv and choose different seed to see if the cv varies a lot. Some of them are high consistently. For example,</p>\n\n<pre><code>('Train drivers: ', ['p022', 'p049', 'p021', 'p026', 'p002', 'p041', 'p042', 'p045', 'p047', 'p024', 'p016', 'p014', 'p039', 'p012', 'p035', 'p052', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\n('Test drivers: ', ['p072', 'p075', 'p015', 'p051'])\n</code></pre>\n\n<p>The reason could be that the test drivers contain p072 which is a large lady and hard to classify.</p>",
      "rawMarkdown": "Lucian, the lb is higher than the local cv.\r\n\r\n@Neil, I saw something like this too. Usually, I do 5-10 folder cv and choose different seed to see if the cv varies a lot. Some of them are high consistently. For example,\r\n\r\n    ('Train drivers: ', ['p022', 'p049', 'p021', 'p026', 'p002', 'p041', 'p042', 'p045', 'p047', 'p024', 'p016', 'p014', 'p039', 'p012', 'p035', 'p052', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\r\n    ('Test drivers: ', ['p072', 'p075', 'p015', 'p051'])\r\n\r\nThe reason could be that the test drivers contain p072 which is a large lady and hard to classify.",
      "votes": null
    },
    {
      "id": "117735",
      "postDate": "04/30/2016 16:09:41",
      "content": "<p>This split is bad too.</p>\n\n<pre><code>('Train drivers: ', ['p022', 'p021', 'p026', 'p002', 'p042', 'p045', 'p047', 'p024', 'p072', 'p075', 'p016', 'p015', 'p014', 'p012', 'p035', 'p052', 'p051', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\n('Test drivers: ', ['p049', 'p041', 'p039'])\n</code></pre>\n\n<p>The training and testing error never go below 2.30.</p>",
      "rawMarkdown": "This split is bad too.\r\n\r\n    ('Train drivers: ', ['p022', 'p021', 'p026', 'p002', 'p042', 'p045', 'p047', 'p024', 'p072', 'p075', 'p016', 'p015', 'p014', 'p012', 'p035', 'p052', 'p051', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\r\n    ('Test drivers: ', ['p049', 'p041', 'p039'])\r\n\r\n  The training and testing error never go below 2.30.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117666,
      "author_name": "ilucian",
      "author_url": "",
      "post_date": "04/30/2016 04:32:06",
      "content": "<p>You might get lucky to find a set of parameters that happens to fit the distribution of the test data really well. I've seen it when using aggressive regularization. Does it happen consistently?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117668,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "04/30/2016 04:40:02",
      "content": "<p>@Lucian, actually not. For specific parameters, it is. But the thing is, the score of lb is not as expected. I don't know what happened. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117669,
      "author_name": "ilucian",
      "author_url": "",
      "post_date": "04/30/2016 04:50:42",
      "content": "<p>Not as expected as in too high or too low?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117676,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/30/2016 06:21:24",
      "content": "<p>When splitting cv by driver, I have seen lower estimated loss for certain drivers consistently. Drivers<code>p045</code> and <code>p051</code> will consistently score lower than training loss, and it seems quite hard to overfit them just by using dropout for regularisation.</p>\n\n<p>One thing to check if it varies - are you always splitting your CV set the exact same way - i.e. same data in it each time (hopefully it is also split by driver)? Taking a different CV set each time could be creating this effect artificially for you since each driver can score very differently.</p>\n\n<p>I tracked this effect in some of my training. Here, the training loss is around 0.7 each time (not recorded, sorry), but the CV loss is known for each driver because I did 26-fold CV:</p>\n\n<pre><code>'p026' =&gt; 1.27, 1.65, 1.53\n'p047' =&gt; 1.33, 1.17, 1.07\n'p045' =&gt; 0.58, 0.51, 0.66\n'p022' =&gt; 1.42, 1.67, 1.31\n'p035' =&gt; 0.83, 0.95, 1.08\n'p024' =&gt; 1.21, 1.35, 1.45\n'p015' =&gt; 2.13, 2.59, 2.51\n'p061' =&gt; 1.21, 1.23, 1.27\n'p056' =&gt; 1.10, 1.03, 1.12\n'p021' =&gt; 1.02, 1.09, 1.19\n'p081' =&gt; 2.07, 1.88, 1.95\n'p039' =&gt; 1.51, 2.26, 1.56\n'p014' =&gt; 1.71, 1.78, 1.46\n'p075' =&gt; 2.00, 2.03, 1.96\n'p002' =&gt; 1.57, 1.72, 1.94\n'p049' =&gt; 1.12, 1.18, 1.24\n'p041' =&gt; 0.91, 0.86, 0.68\n'p016' =&gt; 1.21, 1.24, 1.42\n'p012' =&gt; 1.24, 1.41, 1.41\n'p066' =&gt; 1.70, 1.82, 1.69\n'p042' =&gt; 0.61, 0.63, 1.13\n'p051' =&gt; 0.53, 0.49, 0.56\n'p064' =&gt; 1.68, 1.52, 1.48\n'p052' =&gt; 0.88, 0.85, 0.71\n'p072' =&gt; 2.63, 6.35, 2.37\n'p050' =&gt; 2.27, 1.88, 1.93\n</code></pre>\n\n<p>The columns correspond to slight changes in architecture, and score 1.13, 1.14 and 1.17 respectively on the public leaderboard.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117712,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "04/30/2016 13:50:25",
      "content": "<p>@Lucian, the lb is higher than the local cv.</p>\n\n<p>@Neil, I saw something like this too. Usually, I do 5-10 folder cv and choose different seed to see if the cv varies a lot. Some of them are high consistently. For example,</p>\n\n<pre><code>('Train drivers: ', ['p022', 'p049', 'p021', 'p026', 'p002', 'p041', 'p042', 'p045', 'p047', 'p024', 'p016', 'p014', 'p039', 'p012', 'p035', 'p052', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\n('Test drivers: ', ['p072', 'p075', 'p015', 'p051'])\n</code></pre>\n\n<p>The reason could be that the test drivers contain p072 which is a large lady and hard to classify.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117735,
      "author_name": "zhugds",
      "author_url": "",
      "post_date": "04/30/2016 16:09:41",
      "content": "<p>This split is bad too.</p>\n\n<pre><code>('Train drivers: ', ['p022', 'p021', 'p026', 'p002', 'p042', 'p045', 'p047', 'p024', 'p072', 'p075', 'p016', 'p015', 'p014', 'p012', 'p035', 'p052', 'p051', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\n('Test drivers: ', ['p049', 'p041', 'p039'])\n</code></pre>\n\n<p>The training and testing error never go below 2.30.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117665": "I sometimes encountered this when trying different parameters. Is this normal? What are the possible reasons leading to this? Anyone has any idea?",
    "117666": "You might get lucky to find a set of parameters that happens to fit the distribution of the test data really well. I've seen it when using aggressive regularization. Does it happen consistently?",
    "117668": "Lucian, actually not. For specific parameters, it is. But the thing is, the score of lb is not as expected. I don't know what happened.",
    "117669": "Not as expected as in too high or too low?",
    "117676": "When splitting cv by driver, I have seen lower estimated loss for certain drivers consistently. Drivers`p045` and `p051` will consistently score lower than training loss, and it seems quite hard to overfit them just by using dropout for regularisation.\r\n\r\nOne thing to check if it varies - are you always splitting your CV set the exact same way - i.e. same data in it each time (hopefully it is also split by driver)? Taking a different CV set each time could be creating this effect artificially for you since each driver can score very differently.\r\n\r\nI tracked this effect in some of my training. Here, the training loss is around 0.7 each time (not recorded, sorry), but the CV loss is known for each driver because I did 26-fold CV:\r\n    \r\n    'p026' => 1.27, 1.65, 1.53\r\n    'p047' => 1.33, 1.17, 1.07\r\n    'p045' => 0.58, 0.51, 0.66\r\n    'p022' => 1.42, 1.67, 1.31\r\n    'p035' => 0.83, 0.95, 1.08\r\n    'p024' => 1.21, 1.35, 1.45\r\n    'p015' => 2.13, 2.59, 2.51\r\n    'p061' => 1.21, 1.23, 1.27\r\n    'p056' => 1.10, 1.03, 1.12\r\n    'p021' => 1.02, 1.09, 1.19\r\n    'p081' => 2.07, 1.88, 1.95\r\n    'p039' => 1.51, 2.26, 1.56\r\n    'p014' => 1.71, 1.78, 1.46\r\n    'p075' => 2.00, 2.03, 1.96\r\n    'p002' => 1.57, 1.72, 1.94\r\n    'p049' => 1.12, 1.18, 1.24\r\n    'p041' => 0.91, 0.86, 0.68\r\n    'p016' => 1.21, 1.24, 1.42\r\n    'p012' => 1.24, 1.41, 1.41\r\n    'p066' => 1.70, 1.82, 1.69\r\n    'p042' => 0.61, 0.63, 1.13\r\n    'p051' => 0.53, 0.49, 0.56\r\n    'p064' => 1.68, 1.52, 1.48\r\n    'p052' => 0.88, 0.85, 0.71\r\n    'p072' => 2.63, 6.35, 2.37\r\n    'p050' => 2.27, 1.88, 1.93\r\n\r\nThe columns correspond to slight changes in architecture, and score 1.13, 1.14 and 1.17 respectively on the public leaderboard.",
    "117712": "Lucian, the lb is higher than the local cv.\r\n\r\n@Neil, I saw something like this too. Usually, I do 5-10 folder cv and choose different seed to see if the cv varies a lot. Some of them are high consistently. For example,\r\n\r\n    ('Train drivers: ', ['p022', 'p049', 'p021', 'p026', 'p002', 'p041', 'p042', 'p045', 'p047', 'p024', 'p016', 'p014', 'p039', 'p012', 'p035', 'p052', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\r\n    ('Test drivers: ', ['p072', 'p075', 'p015', 'p051'])\r\n\r\nThe reason could be that the test drivers contain p072 which is a large lady and hard to classify.",
    "117735": "This split is bad too.\r\n\r\n    ('Train drivers: ', ['p022', 'p021', 'p026', 'p002', 'p042', 'p045', 'p047', 'p024', 'p072', 'p075', 'p016', 'p015', 'p014', 'p012', 'p035', 'p052', 'p051', 'p050', 'p056', 'p066', 'p064', 'p061', 'p081'])\r\n    ('Test drivers: ', ['p049', 'p041', 'p039'])\r\n\r\n  The training and testing error never go below 2.30."
  },
  "source": "meta"
}