{
  "id": 196640,
  "title": "BIG LB Score Fluctuations",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/196640",
  "author_name": "",
  "post_date": "2020-11-12T04:24:02.106255200Z",
  "votes": 7,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I submitted results from two checkpoints 2000 steps apart from each other (batch size 384), but weirdly one scored ~25 while the other scored ~18.</p>\n<p>Has anyone encountered similar situations? Should I worry about whether it's overfitting the public test set?</p>",
  "messages": [
    {
      "id": "1075891",
      "postDate": "11/12/2020 04:24:02",
      "content": "<p>I submitted results from two checkpoints 2000 steps apart from each other (batch size 384), but weirdly one scored ~25 while the other scored ~18.</p>\n<p>Has anyone encountered similar situations? Should I worry about whether it's overfitting the public test set?</p>",
      "rawMarkdown": "I submitted results from two checkpoints 2000 steps apart from each other (batch size 384), but weirdly one scored ~25 while the other scored ~18.\n\nHas anyone encountered similar situations? Should I worry about whether it's overfitting the public test set?",
      "votes": null
    },
    {
      "id": "1076043",
      "postDate": "11/12/2020 07:54:54",
      "content": "<p>Faced a similar situation.  I got an LB score 21.xxx  with 16M samples. But unfortunately, I got LB score 80.xxx with a total samples 22M (I used batch size 64). But model training loss was 16.xxx . It look like my model is getting overfitting on the training set.</p>",
      "rawMarkdown": "Faced a similar situation.  I got an LB score 21.xxx  with 16M samples. But unfortunately, I got LB score 80.xxx with a total samples 22M (I used batch size 64). But model training loss was 16.xxx . It look like my model is getting overfitting on the training set.",
      "votes": null
    },
    {
      "id": "1076727",
      "postDate": "11/12/2020 19:37:35",
      "content": "<p>Did you use train_full data?</p>",
      "rawMarkdown": "Did you use train_full data?",
      "votes": null
    },
    {
      "id": "1076919",
      "postDate": "11/13/2020 03:25:48",
      "content": "<p>I didn't use train_full </p>",
      "rawMarkdown": "I didn't use train_full",
      "votes": null
    },
    {
      "id": "1076940",
      "postDate": "11/13/2020 04:05:27",
      "content": "<p>Yeah, sounds like you overfit. You should do your local validation via the chopped data set first to see that.</p>",
      "rawMarkdown": "Yeah, sounds like you overfit. You should do your local validation via the chopped data set first to see that.",
      "votes": null
    },
    {
      "id": "1084470",
      "postDate": "11/20/2020 04:15:08",
      "content": "<p>Started using train_full. But I still didn't see better results. completed training on 25M samples from train_full. Got training loss - 19.xxx but LB score is 32.xxx  :(  Am I missing anything here? </p>",
      "rawMarkdown": "Started using train_full. But I still didn't see better results. completed training on 25M samples from train_full. Got training loss - 19.xxx but LB score is 32.xxx  :(  Am I missing anything here?",
      "votes": null
    },
    {
      "id": "1084480",
      "postDate": "11/20/2020 04:26:30",
      "content": "<p>For me, I do see a gap between validation score and train. For the LB score 23.xxx, the corresponding train score is 11.9xx. For train loss 19.xxx, it is reasonable to get 32.xxx LB score. <br>\nAgain, have you test your model on the chopped validation set? Is it similar to LB? </p>",
      "rawMarkdown": "For me, I do see a gap between validation score and train. For the LB score 23.xxx, the corresponding train score is 11.9xx. For train loss 19.xxx, it is reasonable to get 32.xxx LB score. \nAgain, have you test your model on the chopped validation set? Is it similar to LB?",
      "votes": null
    },
    {
      "id": "1084565",
      "postDate": "11/20/2020 06:51:02",
      "content": "<p><strong>For train loss 19.xxx, it is reasonable to get 32.xxx LB score.</strong>  <br>\nBut I got 22.xxx training loss and 26.xxx LB score at 200000 steps with batch size 64.  But why LB score become 32.xxx at 400000 steps with batch size 64?<br>\nFYI : didn't use the chopped validation set.</p>",
      "rawMarkdown": "**For train loss 19.xxx, it is reasonable to get 32.xxx LB score.**  \nBut I got 22.xxx training loss and 26.xxx LB score at 200000 steps with batch size 64.  But why LB score become 32.xxx at 400000 steps with batch size 64?\nFYI : didn't use the chopped validation set.",
      "votes": null
    },
    {
      "id": "1084620",
      "postDate": "11/20/2020 08:01:03",
      "content": "<p>Yeah, then it sounds like it overfit somehow.</p>",
      "rawMarkdown": "Yeah, then it sounds like it overfit somehow.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1076043,
      "author_name": "pallaviroyal",
      "author_url": "",
      "post_date": "11/12/2020 07:54:54",
      "content": "<p>Faced a similar situation.  I got an LB score 21.xxx  with 16M samples. But unfortunately, I got LB score 80.xxx with a total samples 22M (I used batch size 64). But model training loss was 16.xxx . It look like my model is getting overfitting on the training set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1076727,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "11/12/2020 19:37:35",
          "content": "<p>Did you use train_full data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076919,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "11/13/2020 03:25:48",
          "content": "<p>I didn't use train_full </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1076940,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/13/2020 04:05:27",
          "content": "<p>Yeah, sounds like you overfit. You should do your local validation via the chopped data set first to see that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084470,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "11/20/2020 04:15:08",
          "content": "<p>Started using train_full. But I still didn't see better results. completed training on 25M samples from train_full. Got training loss - 19.xxx but LB score is 32.xxx  :(  Am I missing anything here? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084480,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/20/2020 04:26:30",
          "content": "<p>For me, I do see a gap between validation score and train. For the LB score 23.xxx, the corresponding train score is 11.9xx. For train loss 19.xxx, it is reasonable to get 32.xxx LB score. <br>\nAgain, have you test your model on the chopped validation set? Is it similar to LB? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084565,
          "author_name": "pallaviroyal",
          "author_url": "",
          "post_date": "11/20/2020 06:51:02",
          "content": "<p><strong>For train loss 19.xxx, it is reasonable to get 32.xxx LB score.</strong>  <br>\nBut I got 22.xxx training loss and 26.xxx LB score at 200000 steps with batch size 64.  But why LB score become 32.xxx at 400000 steps with batch size 64?<br>\nFYI : didn't use the chopped validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1084620,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "11/20/2020 08:01:03",
          "content": "<p>Yeah, then it sounds like it overfit somehow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1075891": "I submitted results from two checkpoints 2000 steps apart from each other (batch size 384), but weirdly one scored ~25 while the other scored ~18.\n\nHas anyone encountered similar situations? Should I worry about whether it's overfitting the public test set?",
    "1076043": "Faced a similar situation.  I got an LB score 21.xxx  with 16M samples. But unfortunately, I got LB score 80.xxx with a total samples 22M (I used batch size 64). But model training loss was 16.xxx . It look like my model is getting overfitting on the training set.",
    "1076727": "Did you use train_full data?",
    "1076919": "I didn't use train_full",
    "1076940": "Yeah, sounds like you overfit. You should do your local validation via the chopped data set first to see that.",
    "1084470": "Started using train_full. But I still didn't see better results. completed training on 25M samples from train_full. Got training loss - 19.xxx but LB score is 32.xxx  :(  Am I missing anything here?",
    "1084480": "For me, I do see a gap between validation score and train. For the LB score 23.xxx, the corresponding train score is 11.9xx. For train loss 19.xxx, it is reasonable to get 32.xxx LB score. \nAgain, have you test your model on the chopped validation set? Is it similar to LB?",
    "1084565": "**For train loss 19.xxx, it is reasonable to get 32.xxx LB score.**  \nBut I got 22.xxx training loss and 26.xxx LB score at 200000 steps with batch size 64.  But why LB score become 32.xxx at 400000 steps with batch size 64?\nFYI : didn't use the chopped validation set.",
    "1084620": "Yeah, then it sounds like it overfit somehow."
  },
  "source": "meta"
}