{
  "id": 117106,
  "title": "LB score dropped 30% when trained on trainval set",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/117106",
  "author_name": "",
  "post_date": "2019-11-13T12:27:22.206141300Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>We used a train/val split to tune model's performance and uploaded scores to the LB to check that val scores correspond to what we were getting on LB. In the last couple of days before the deadline, we trained the best model using both train and val and got lower scores than training only on train set.  </p>\n\n<p>Did any one else observe a similar behaviour? My only explanation is that the data seems very messy and using more data leads to higher training confusion. </p>\n\n<p>EDIT: I should mention that all hyper parameters were left unchanged  </p>",
  "messages": [
    {
      "id": "671989",
      "postDate": "11/13/2019 12:27:22",
      "content": "<p>We used a train/val split to tune model's performance and uploaded scores to the LB to check that val scores correspond to what we were getting on LB. In the last couple of days before the deadline, we trained the best model using both train and val and got lower scores than training only on train set.  </p>\n\n<p>Did any one else observe a similar behaviour? My only explanation is that the data seems very messy and using more data leads to higher training confusion. </p>\n\n<p>EDIT: I should mention that all hyper parameters were left unchanged  </p>",
      "rawMarkdown": "We used a train/val split to tune model's performance and uploaded scores to the LB to check that val scores correspond to what we were getting on LB. In the last couple of days before the deadline, we trained the best model using both train and val and got lower scores than training only on train set.  \n\nDid any one else observe a similar behaviour? My only explanation is that the data seems very messy and using more data leads to higher training confusion. \n\nEDIT: I should mention that all hyper parameters were left unchanged",
      "votes": null
    },
    {
      "id": "672152",
      "postDate": "11/13/2019 15:05:32",
      "content": "<p>If you trained on more data with the same parameters your training loss could have been higher. I would compare the loss for each model and see if they are close to each other. </p>",
      "rawMarkdown": "If you trained on more data with the same parameters your training loss could have been higher. I would compare the loss for each model and see if they are close to each other.",
      "votes": null
    },
    {
      "id": "672438",
      "postDate": "11/13/2019 21:56:00",
      "content": "<p>I have achieved 0.25 on evaluation split but only 0.15 on publish LB.</p>",
      "rawMarkdown": "I have achieved 0.25 on evaluation split but only 0.15 on publish LB.",
      "votes": null
    },
    {
      "id": "672472",
      "postDate": "11/13/2019 23:41:25",
      "content": "<p>In the training set, there are a few scenes with multiple lidars, but none in the test split.  So perhaps the cars are a little different.  Haven't looked at the timestamps, but the test split might also be from a different month... the drives span 5 months I think?</p>\n\n<p>Did you create your own validation holdout split?  I couldn't find one and chose about 20% scenes randomly.  NuScenes has a designated val split though.</p>",
      "rawMarkdown": "In the training set, there are a few scenes with multiple lidars, but none in the test split.  So perhaps the cars are a little different.  Haven't looked at the timestamps, but the test split might also be from a different month... the drives span 5 months I think?\n\nDid you create your own validation holdout split?  I couldn't find one and chose about 20% scenes randomly.  NuScenes has a designated val split though.",
      "votes": null
    },
    {
      "id": "673123",
      "postDate": "11/14/2019 14:48:03",
      "content": "<p>Thanks for your suggestion, the loss converged to a lower value though. Below is a screenshot with the loss metrics for train only (blue) and trainval (red-ish)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F996734%2F1688af2d06a497880c52fde3e560ff7b%2FScreenshot%202019-11-14%20at%2017.46.33.png?generation=1573742876823859&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for your suggestion, the loss converged to a lower value though. Below is a screenshot with the loss metrics for train only (blue) and trainval (red-ish)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F996734%2F1688af2d06a497880c52fde3e560ff7b%2FScreenshot%202019-11-14%20at%2017.46.33.png?generation=1573742876823859&amp;alt=media)",
      "votes": null
    },
    {
      "id": "673133",
      "postDate": "11/14/2019 14:56:56",
      "content": "<blockquote>\n  <p>In the training set, there are a few scenes with multiple lidars, but none in the test split. So perhaps the cars are a little different. Haven't looked at the timestamps, but the test split might also be from a different month… the drives span 5 months I think?</p>\n</blockquote>\n\n<p>Didn't know about that, perhaps we used the driving sequences from the same timeframe, will check when I have some time</p>\n\n<blockquote>\n  <p>Did you create your own validation holdout split?</p>\n</blockquote>\n\n<p>We actually used a sample provided in the reference model. Maybe shouldve given that a second thought before using something like that though :D</p>",
      "rawMarkdown": "&gt; In the training set, there are a few scenes with multiple lidars, but none in the test split. So perhaps the cars are a little different. Haven't looked at the timestamps, but the test split might also be from a different month… the drives span 5 months I think?\n\nDidn't know about that, perhaps we used the driving sequences from the same timeframe, will check when I have some time\n\n&gt; Did you create your own validation holdout split?\n\nWe actually used a sample provided in the reference model. Maybe shouldve given that a second thought before using something like that though :D",
      "votes": null
    },
    {
      "id": "673134",
      "postDate": "11/14/2019 14:57:46",
      "content": "<p>In out case we were getting very similar results on val and public LB, but when we trained on both train and val, we got -30% on public LB :( </p>",
      "rawMarkdown": "In out case we were getting very similar results on val and public LB, but when we trained on both train and val, we got -30% on public LB :(",
      "votes": null
    },
    {
      "id": "673391",
      "postDate": "11/14/2019 23:13:20",
      "content": "<p>Looks like more data more better!  Do you happen to have the validation set loss to compare?  If the model is good, then val loss should probably be close to trainval loss, and hopefully val loss is not higher than train (blue) loss.  (Overfitting).</p>\n\n<p>I'm still curious that there might be some adjustment that you were applying in your own evaluation that didn't get applied server-side (or with the test set).  Do you know how you were transforming predictions from the lidar frame to the world frame?  The NuScenes code may transform / interpolate point clouds and/or cuboids based upon their timestamp.  If the Kaggle server and your own code disagree, this could turn what look like good predictions into terrible predictions.  </p>",
      "rawMarkdown": "Looks like more data more better!  Do you happen to have the validation set loss to compare?  If the model is good, then val loss should probably be close to trainval loss, and hopefully val loss is not higher than train (blue) loss.  (Overfitting).\n\nI'm still curious that there might be some adjustment that you were applying in your own evaluation that didn't get applied server-side (or with the test set).  Do you know how you were transforming predictions from the lidar frame to the world frame?  The NuScenes code may transform / interpolate point clouds and/or cuboids based upon their timestamp.  If the Kaggle server and your own code disagree, this could turn what look like good predictions into terrible predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 672152,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "11/13/2019 15:05:32",
      "content": "<p>If you trained on more data with the same parameters your training loss could have been higher. I would compare the loss for each model and see if they are close to each other. </p>",
      "votes": null,
      "replies": [
        {
          "id": 673123,
          "author_name": "ruslanm",
          "author_url": "",
          "post_date": "11/14/2019 14:48:03",
          "content": "<p>Thanks for your suggestion, the loss converged to a lower value though. Below is a screenshot with the loss metrics for train only (blue) and trainval (red-ish)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F996734%2F1688af2d06a497880c52fde3e560ff7b%2FScreenshot%202019-11-14%20at%2017.46.33.png?generation=1573742876823859&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673391,
          "author_name": "oarphme",
          "author_url": "",
          "post_date": "11/14/2019 23:13:20",
          "content": "<p>Looks like more data more better!  Do you happen to have the validation set loss to compare?  If the model is good, then val loss should probably be close to trainval loss, and hopefully val loss is not higher than train (blue) loss.  (Overfitting).</p>\n\n<p>I'm still curious that there might be some adjustment that you were applying in your own evaluation that didn't get applied server-side (or with the test set).  Do you know how you were transforming predictions from the lidar frame to the world frame?  The NuScenes code may transform / interpolate point clouds and/or cuboids based upon their timestamp.  If the Kaggle server and your own code disagree, this could turn what look like good predictions into terrible predictions.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 672438,
      "author_name": "hanxiaodeng",
      "author_url": "",
      "post_date": "11/13/2019 21:56:00",
      "content": "<p>I have achieved 0.25 on evaluation split but only 0.15 on publish LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 673134,
          "author_name": "ruslanm",
          "author_url": "",
          "post_date": "11/14/2019 14:57:46",
          "content": "<p>In out case we were getting very similar results on val and public LB, but when we trained on both train and val, we got -30% on public LB :( </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 672472,
      "author_name": "oarphme",
      "author_url": "",
      "post_date": "11/13/2019 23:41:25",
      "content": "<p>In the training set, there are a few scenes with multiple lidars, but none in the test split.  So perhaps the cars are a little different.  Haven't looked at the timestamps, but the test split might also be from a different month... the drives span 5 months I think?</p>\n\n<p>Did you create your own validation holdout split?  I couldn't find one and chose about 20% scenes randomly.  NuScenes has a designated val split though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 673133,
          "author_name": "ruslanm",
          "author_url": "",
          "post_date": "11/14/2019 14:56:56",
          "content": "<blockquote>\n  <p>In the training set, there are a few scenes with multiple lidars, but none in the test split. So perhaps the cars are a little different. Haven't looked at the timestamps, but the test split might also be from a different month… the drives span 5 months I think?</p>\n</blockquote>\n\n<p>Didn't know about that, perhaps we used the driving sequences from the same timeframe, will check when I have some time</p>\n\n<blockquote>\n  <p>Did you create your own validation holdout split?</p>\n</blockquote>\n\n<p>We actually used a sample provided in the reference model. Maybe shouldve given that a second thought before using something like that though :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "671989": "We used a train/val split to tune model's performance and uploaded scores to the LB to check that val scores correspond to what we were getting on LB. In the last couple of days before the deadline, we trained the best model using both train and val and got lower scores than training only on train set.  \n\nDid any one else observe a similar behaviour? My only explanation is that the data seems very messy and using more data leads to higher training confusion. \n\nEDIT: I should mention that all hyper parameters were left unchanged",
    "672152": "If you trained on more data with the same parameters your training loss could have been higher. I would compare the loss for each model and see if they are close to each other.",
    "672438": "I have achieved 0.25 on evaluation split but only 0.15 on publish LB.",
    "672472": "In the training set, there are a few scenes with multiple lidars, but none in the test split.  So perhaps the cars are a little different.  Haven't looked at the timestamps, but the test split might also be from a different month... the drives span 5 months I think?\n\nDid you create your own validation holdout split?  I couldn't find one and chose about 20% scenes randomly.  NuScenes has a designated val split though.",
    "673123": "Thanks for your suggestion, the loss converged to a lower value though. Below is a screenshot with the loss metrics for train only (blue) and trainval (red-ish)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F996734%2F1688af2d06a497880c52fde3e560ff7b%2FScreenshot%202019-11-14%20at%2017.46.33.png?generation=1573742876823859&amp;alt=media)",
    "673133": "&gt; In the training set, there are a few scenes with multiple lidars, but none in the test split. So perhaps the cars are a little different. Haven't looked at the timestamps, but the test split might also be from a different month… the drives span 5 months I think?\n\nDidn't know about that, perhaps we used the driving sequences from the same timeframe, will check when I have some time\n\n&gt; Did you create your own validation holdout split?\n\nWe actually used a sample provided in the reference model. Maybe shouldve given that a second thought before using something like that though :D",
    "673134": "In out case we were getting very similar results on val and public LB, but when we trained on both train and val, we got -30% on public LB :(",
    "673391": "Looks like more data more better!  Do you happen to have the validation set loss to compare?  If the model is good, then val loss should probably be close to trainval loss, and hopefully val loss is not higher than train (blue) loss.  (Overfitting).\n\nI'm still curious that there might be some adjustment that you were applying in your own evaluation that didn't get applied server-side (or with the test set).  Do you know how you were transforming predictions from the lidar frame to the world frame?  The NuScenes code may transform / interpolate point clouds and/or cuboids based upon their timestamp.  If the Kaggle server and your own code disagree, this could turn what look like good predictions into terrible predictions."
  },
  "source": "meta"
}