{
  "id": 55498,
  "title": "the val score is always higher than train score in training time",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55498",
  "author_name": "",
  "post_date": "2018-04-27T09:17:16.459845700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I use the last 80M to do model train and the last 10M as validation data, but the val score is higher than train score in mostly time,  Why?  I  think maybe the val data size is not enough , and in this case the val data is not very suitable to judge the model, and I use the val score to do early stop, is it too early because the val data is not well. Did your val score is also like as mine ? Any help is apprecitaed!</p>\n\n<p>Such like this :</p>\n\n<p>Training until validation scores don't improve for 30 rounds.\n[10]    train's auc: 0.968284   valid's auc: 0.970434\n[20]    train's auc: 0.973109   valid's auc: 0.97454\n[30]    train's auc: 0.974275   valid's auc: 0.975837\n[40]    train's auc: 0.975717   valid's auc: 0.977616\n[50]    train's auc: 0.976816   valid's auc: 0.978618\n[60]    train's auc: 0.977915   valid's auc: 0.980098\n[70]    train's auc: 0.978943   valid's auc: 0.981319\n[80]    train's auc: 0.979702   valid's auc: 0.982267\n[90]    train's auc: 0.980317   valid's auc: 0.983178\n[100]   train's auc: 0.980832   valid's auc: 0.983897\n[110]   train's auc: 0.981192   valid's auc: 0.984378\n[120]   train's auc: 0.981509   valid's auc: 0.984746\n[130]   train's auc: 0.981751   valid's auc: 0.985034</p>",
  "messages": [
    {
      "id": "320006",
      "postDate": "04/27/2018 09:17:16",
      "content": "<p>I use the last 80M to do model train and the last 10M as validation data, but the val score is higher than train score in mostly time,  Why?  I  think maybe the val data size is not enough , and in this case the val data is not very suitable to judge the model, and I use the val score to do early stop, is it too early because the val data is not well. Did your val score is also like as mine ? Any help is apprecitaed!</p>\n\n<p>Such like this :</p>\n\n<p>Training until validation scores don't improve for 30 rounds.\n[10]    train's auc: 0.968284   valid's auc: 0.970434\n[20]    train's auc: 0.973109   valid's auc: 0.97454\n[30]    train's auc: 0.974275   valid's auc: 0.975837\n[40]    train's auc: 0.975717   valid's auc: 0.977616\n[50]    train's auc: 0.976816   valid's auc: 0.978618\n[60]    train's auc: 0.977915   valid's auc: 0.980098\n[70]    train's auc: 0.978943   valid's auc: 0.981319\n[80]    train's auc: 0.979702   valid's auc: 0.982267\n[90]    train's auc: 0.980317   valid's auc: 0.983178\n[100]   train's auc: 0.980832   valid's auc: 0.983897\n[110]   train's auc: 0.981192   valid's auc: 0.984378\n[120]   train's auc: 0.981509   valid's auc: 0.984746\n[130]   train's auc: 0.981751   valid's auc: 0.985034</p>",
      "rawMarkdown": "I use the last 80M to do model train and the last 10M as validation data, but the val score is higher than train score in mostly time,  Why?  I  think maybe the val data size is not enough , and in this case the val data is not very suitable to judge the model, and I use the val score to do early stop, is it too early because the val data is not well. Did your val score is also like as mine ? Any help is apprecitaed!\n\nSuch like this :\n\nTraining until validation scores don't improve for 30 rounds.\n[10]\ttrain's auc: 0.968284\tvalid's auc: 0.970434\n[20]\ttrain's auc: 0.973109\tvalid's auc: 0.97454\n[30]\ttrain's auc: 0.974275\tvalid's auc: 0.975837\n[40]\ttrain's auc: 0.975717\tvalid's auc: 0.977616\n[50]\ttrain's auc: 0.976816\tvalid's auc: 0.978618\n[60]\ttrain's auc: 0.977915\tvalid's auc: 0.980098\n[70]\ttrain's auc: 0.978943\tvalid's auc: 0.981319\n[80]\ttrain's auc: 0.979702\tvalid's auc: 0.982267\n[90]\ttrain's auc: 0.980317\tvalid's auc: 0.983178\n[100]\ttrain's auc: 0.980832\tvalid's auc: 0.983897\n[110]\ttrain's auc: 0.981192\tvalid's auc: 0.984378\n[120]\ttrain's auc: 0.981509\tvalid's auc: 0.984746\n[130]\ttrain's auc: 0.981751\tvalid's auc: 0.985034",
      "votes": null
    },
    {
      "id": "320225",
      "postDate": "04/27/2018 23:12:02",
      "content": "<p>8000w 和 1000w是指 八千万和一千万么？我猜老外可能不太懂w是代表一万，估计你说80M和10M，会更容易让大家理解。</p>\n\n<p>English version: the 'w' in 8000w and 1000w means 10000, so 8000w = 80million and 1000w means 10million. </p>",
      "rawMarkdown": "8000w 和 1000w是指 八千万和一千万么？我猜老外可能不太懂w是代表一万，估计你说80M和10M，会更容易让大家理解。\n\nEnglish version: the 'w' in 8000w and 1000w means 10000, so 8000w = 80million and 1000w means 10million.",
      "votes": null
    },
    {
      "id": "320241",
      "postDate": "04/28/2018 01:31:36",
      "content": "<p>Maybe your model is underfitting data.  what is your model accuracy,precision,recall on trainning data/val data ?\nYou can try to  add more features or select a new model ,I think.</p>",
      "rawMarkdown": "Maybe your model is underfitting data.  what is your model accuracy,precision,recall on trainning data/val data ?\nYou can try to  add more features or select a new model ,I think.",
      "votes": null
    },
    {
      "id": "320299",
      "postDate": "04/28/2018 07:18:22",
      "content": "<p>xiexie</p>",
      "rawMarkdown": "xiexie",
      "votes": null
    },
    {
      "id": "320375",
      "postDate": "04/28/2018 13:40:10",
      "content": "<p>Ohaha， thanks very much</p>",
      "rawMarkdown": "Ohaha， thanks very much",
      "votes": null
    },
    {
      "id": "320419",
      "postDate": "04/28/2018 16:14:08",
      "content": "<p>that's the feature of this dataset. If you use last rows as validation, your validation auc should be above 0.990.</p>",
      "rawMarkdown": "that's the feature of this dataset. If you use last rows as validation, your validation auc should be above 0.990.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 320225,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "04/27/2018 23:12:02",
      "content": "<p>8000w 和 1000w是指 八千万和一千万么？我猜老外可能不太懂w是代表一万，估计你说80M和10M，会更容易让大家理解。</p>\n\n<p>English version: the 'w' in 8000w and 1000w means 10000, so 8000w = 80million and 1000w means 10million. </p>",
      "votes": null,
      "replies": [
        {
          "id": 320299,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "04/28/2018 07:18:22",
          "content": "<p>xiexie</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320375,
          "author_name": "qfzgs1994",
          "author_url": "",
          "post_date": "04/28/2018 13:40:10",
          "content": "<p>Ohaha， thanks very much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 320241,
      "author_name": "chinaxwz",
      "author_url": "",
      "post_date": "04/28/2018 01:31:36",
      "content": "<p>Maybe your model is underfitting data.  what is your model accuracy,precision,recall on trainning data/val data ?\nYou can try to  add more features or select a new model ,I think.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320419,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "04/28/2018 16:14:08",
      "content": "<p>that's the feature of this dataset. If you use last rows as validation, your validation auc should be above 0.990.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "320006": "I use the last 80M to do model train and the last 10M as validation data, but the val score is higher than train score in mostly time,  Why?  I  think maybe the val data size is not enough , and in this case the val data is not very suitable to judge the model, and I use the val score to do early stop, is it too early because the val data is not well. Did your val score is also like as mine ? Any help is apprecitaed!\n\nSuch like this :\n\nTraining until validation scores don't improve for 30 rounds.\n[10]\ttrain's auc: 0.968284\tvalid's auc: 0.970434\n[20]\ttrain's auc: 0.973109\tvalid's auc: 0.97454\n[30]\ttrain's auc: 0.974275\tvalid's auc: 0.975837\n[40]\ttrain's auc: 0.975717\tvalid's auc: 0.977616\n[50]\ttrain's auc: 0.976816\tvalid's auc: 0.978618\n[60]\ttrain's auc: 0.977915\tvalid's auc: 0.980098\n[70]\ttrain's auc: 0.978943\tvalid's auc: 0.981319\n[80]\ttrain's auc: 0.979702\tvalid's auc: 0.982267\n[90]\ttrain's auc: 0.980317\tvalid's auc: 0.983178\n[100]\ttrain's auc: 0.980832\tvalid's auc: 0.983897\n[110]\ttrain's auc: 0.981192\tvalid's auc: 0.984378\n[120]\ttrain's auc: 0.981509\tvalid's auc: 0.984746\n[130]\ttrain's auc: 0.981751\tvalid's auc: 0.985034",
    "320225": "8000w 和 1000w是指 八千万和一千万么？我猜老外可能不太懂w是代表一万，估计你说80M和10M，会更容易让大家理解。\n\nEnglish version: the 'w' in 8000w and 1000w means 10000, so 8000w = 80million and 1000w means 10million.",
    "320241": "Maybe your model is underfitting data.  what is your model accuracy,precision,recall on trainning data/val data ?\nYou can try to  add more features or select a new model ,I think.",
    "320299": "xiexie",
    "320375": "Ohaha， thanks very much",
    "320419": "that's the feature of this dataset. If you use last rows as validation, your validation auc should be above 0.990."
  },
  "source": "meta"
}