{
  "id": 55984,
  "title": "how to fix great gap between local val score and LB",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55984",
  "author_name": "",
  "post_date": "2018-05-04T02:58:47.134413800Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>100M dataset for training, 5M data for validation,32features, loca val score is 0.98378, LB is 0.9800.There is huge gap.Did someone have the same problem？</p>",
  "messages": [
    {
      "id": "322959",
      "postDate": "05/04/2018 02:58:47",
      "content": "<p>100M dataset for training, 5M data for validation,32features, loca val score is 0.98378, LB is 0.9800.There is huge gap.Did someone have the same problem？</p>",
      "rawMarkdown": "100M dataset for training, 5M data for validation,32features, loca val score is 0.98378, LB is 0.9800.There is huge gap.Did someone have the same problem？",
      "votes": null
    },
    {
      "id": "322961",
      "postDate": "05/04/2018 03:03:48",
      "content": "<p>try 20M validation and see if the gap shrink a bit!</p>",
      "rawMarkdown": "try 20M validation and see if the gap shrink a bit!",
      "votes": null
    },
    {
      "id": "322962",
      "postDate": "05/04/2018 03:06:31",
      "content": "<p>validation was random split</p>",
      "rawMarkdown": "validation was random split",
      "votes": null
    },
    {
      "id": "322964",
      "postDate": "05/04/2018 03:08:41",
      "content": "<p>that maybe not work, for I uses 10m for val ,val score went up and LB score went down</p>",
      "rawMarkdown": "that maybe not work, for I uses 10m for val ,val score went up and LB score went down",
      "votes": null
    },
    {
      "id": "323138",
      "postDate": "05/04/2018 13:06:01",
      "content": "<p>Random split validation doesn't necessarily make sense in a setting where there are meaningful time-based differences in the data. Instead, you can get a very reliable validation setting (for the public LB at least) by mirroring the time ranges in the test data. Personally I use the test hours of day 9 for validation.</p>",
      "rawMarkdown": "Random split validation doesn't necessarily make sense in a setting where there are meaningful time-based differences in the data. Instead, you can get a very reliable validation setting (for the public LB at least) by mirroring the time ranges in the test data. Personally I use the test hours of day 9 for validation.",
      "votes": null
    },
    {
      "id": "323185",
      "postDate": "05/04/2018 15:02:56",
      "content": "<p>good idea! I will try it, thx</p>",
      "rawMarkdown": "good idea! I will try it, thx",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 322961,
      "author_name": "danieleewww",
      "author_url": "",
      "post_date": "05/04/2018 03:03:48",
      "content": "<p>try 20M validation and see if the gap shrink a bit!</p>",
      "votes": null,
      "replies": [
        {
          "id": 322964,
          "author_name": "cnzjhdx",
          "author_url": "",
          "post_date": "05/04/2018 03:08:41",
          "content": "<p>that maybe not work, for I uses 10m for val ,val score went up and LB score went down</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 322962,
      "author_name": "cnzjhdx",
      "author_url": "",
      "post_date": "05/04/2018 03:06:31",
      "content": "<p>validation was random split</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323138,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/04/2018 13:06:01",
      "content": "<p>Random split validation doesn't necessarily make sense in a setting where there are meaningful time-based differences in the data. Instead, you can get a very reliable validation setting (for the public LB at least) by mirroring the time ranges in the test data. Personally I use the test hours of day 9 for validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 323185,
          "author_name": "cnzjhdx",
          "author_url": "",
          "post_date": "05/04/2018 15:02:56",
          "content": "<p>good idea! I will try it, thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "322959": "100M dataset for training, 5M data for validation,32features, loca val score is 0.98378, LB is 0.9800.There is huge gap.Did someone have the same problem？",
    "322961": "try 20M validation and see if the gap shrink a bit!",
    "322962": "validation was random split",
    "322964": "that maybe not work, for I uses 10m for val ,val score went up and LB score went down",
    "323138": "Random split validation doesn't necessarily make sense in a setting where there are meaningful time-based differences in the data. Instead, you can get a very reliable validation setting (for the public LB at least) by mirroring the time ranges in the test data. Personally I use the test hours of day 9 for validation.",
    "323185": "good idea! I will try it, thx"
  },
  "source": "meta"
}