{
  "id": 159569,
  "title": "Impact on Training on the Validation Data on LB Score",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159569",
  "author_name": "",
  "post_date": "2020-06-17T20:17:47.353737100Z",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>We know that there are ~600 samples in the LB that are in the validation set.</p>\n\n<p>However, a lot of people are training on the validation set, which will results in over-estimated public LB scores. </p>\n\n<p>I wanted to check by how much approximately, and for a model that scores CV 0.945~, you can expect a boost varying from <strong>0.001 to 0.002</strong> from the leakage... Which is quite a lot considering how stacked the LB is.</p>\n\n<p>So make sure you validate your models correctly ! </p>\n\n<p>EDIT : 600 samples in the complete LB, not public</p>\n\n<p>Code : \n```\nimport numpy as np\nfrom sklearn.metrics import *\nfrom tqdm.notebook import tqdm</p>\n\n<p>base_scores = []\ndeltas = []</p>\n\n<p>truth = np.concatenate((np.ones(7000), np.zeros(56000)))  # 63000 sample\nprint(f'Class 1 ratio : {truth.sum() / truth.size :.3f}') # 0.111</p>\n\n<p>for i in tqdm(range(1000)):\n    pred = (truth / 2 + 0.25) + (np.random.random(size=truth.size) * 0.74 - 0.5) \n    base_score = roc_auc_score(truth, pred)</p>\n\n<pre><code># Leakage : Setting 600 labels to their ground truth\npred[:70] = 0.99\npred[-530:] = 0.01\n\nnew_score = roc_auc_score(truth, pred)\n\nbase_scores.append(base_score)\ndeltas.append(new_score - base_score)\n</code></pre>\n\n<p>print(f'Average score : {np.mean(base_scores):.4f} - Average boost : {np.mean(deltas):.4f}')</p>\n\n<h1>Average score : 0.9474 - Average boost : 0.0010</h1>\n\n<p>```</p>",
  "messages": [
    {
      "id": "890999",
      "postDate": "06/17/2020 20:17:47",
      "content": "<p>We know that there are ~600 samples in the LB that are in the validation set.</p>\n\n<p>However, a lot of people are training on the validation set, which will results in over-estimated public LB scores. </p>\n\n<p>I wanted to check by how much approximately, and for a model that scores CV 0.945~, you can expect a boost varying from <strong>0.001 to 0.002</strong> from the leakage... Which is quite a lot considering how stacked the LB is.</p>\n\n<p>So make sure you validate your models correctly ! </p>\n\n<p>EDIT : 600 samples in the complete LB, not public</p>\n\n<p>Code : \n```\nimport numpy as np\nfrom sklearn.metrics import *\nfrom tqdm.notebook import tqdm</p>\n\n<p>base_scores = []\ndeltas = []</p>\n\n<p>truth = np.concatenate((np.ones(7000), np.zeros(56000)))  # 63000 sample\nprint(f'Class 1 ratio : {truth.sum() / truth.size :.3f}') # 0.111</p>\n\n<p>for i in tqdm(range(1000)):\n    pred = (truth / 2 + 0.25) + (np.random.random(size=truth.size) * 0.74 - 0.5) \n    base_score = roc_auc_score(truth, pred)</p>\n\n<pre><code># Leakage : Setting 600 labels to their ground truth\npred[:70] = 0.99\npred[-530:] = 0.01\n\nnew_score = roc_auc_score(truth, pred)\n\nbase_scores.append(base_score)\ndeltas.append(new_score - base_score)\n</code></pre>\n\n<p>print(f'Average score : {np.mean(base_scores):.4f} - Average boost : {np.mean(deltas):.4f}')</p>\n\n<h1>Average score : 0.9474 - Average boost : 0.0010</h1>\n\n<p>```</p>",
      "rawMarkdown": "We know that there are ~600 samples in the LB that are in the validation set.\n\nHowever, a lot of people are training on the validation set, which will results in over-estimated public LB scores. \n\nI wanted to check by how much approximately, and for a model that scores CV 0.945~, you can expect a boost varying from **0.001 to 0.002** from the leakage... Which is quite a lot considering how stacked the LB is.\n\nSo make sure you validate your models correctly ! \n\nEDIT : 600 samples in the complete LB, not public\n\nCode : \n```\nimport numpy as np\nfrom sklearn.metrics import *\nfrom tqdm.notebook import tqdm\n\nbase_scores = []\ndeltas = []\n\ntruth = np.concatenate((np.ones(7000), np.zeros(56000)))  # 63000 sample\nprint(f'Class 1 ratio : {truth.sum() / truth.size :.3f}') # 0.111\n\nfor i in tqdm(range(1000)):\n    pred = (truth / 2 + 0.25) + (np.random.random(size=truth.size) * 0.74 - 0.5) \n    base_score = roc_auc_score(truth, pred)\n\n    # Leakage : Setting 600 labels to their ground truth\n    pred[:70] = 0.99\n    pred[-530:] = 0.01\n\n    new_score = roc_auc_score(truth, pred)\n\n    base_scores.append(base_score)\n    deltas.append(new_score - base_score)\n\nprint(f'Average score : {np.mean(base_scores):.4f} - Average boost : {np.mean(deltas):.4f}')\n# Average score : 0.9474 - Average boost : 0.0010\n```",
      "votes": null
    },
    {
      "id": "891000",
      "postDate": "06/17/2020 20:18:32",
      "content": "<p>Thanks for the info</p>",
      "rawMarkdown": "Thanks for the info",
      "votes": null
    },
    {
      "id": "891174",
      "postDate": "06/18/2020 01:23:20",
      "content": "<p>Very helpful. Thanks for the info.</p>",
      "rawMarkdown": "Very helpful. Thanks for the info.",
      "votes": null
    },
    {
      "id": "891236",
      "postDate": "06/18/2020 03:26:01",
      "content": "<p>Thank you for sharing <a href=\"/theoviel\">@theoviel</a> . I must have missed something. \n1) How is it discovered that the 600 samples from validation.csv are in the public LB set and not partly in the private test set)?\n2) I thought that if a large model is trained on samples including these 600 then it would mostly classify correctly the copies of all those samples in the test set - so, no improvement in the PL score should be expected. Please correct me.</p>",
      "rawMarkdown": "Thank you for sharing @theoviel . I must have missed something. \n1) How is it discovered that the 600 samples from validation.csv are in the public LB set and not partly in the private test set)?\n2) I thought that if a large model is trained on samples including these 600 then it would mostly classify correctly the copies of all those samples in the test set - so, no improvement in the PL score should be expected. Please correct me.",
      "votes": null
    },
    {
      "id": "891355",
      "postDate": "06/18/2020 05:51:33",
      "content": "<p><a href=\"/isakev\">@isakev</a> there is a little mistake,sorry for that, we mean there are 600 validation samples that  are already in test data(not necessarily in public part) ,,, please check this <a href=\"https://www.kaggle.com/sai11fkaneko/data-leak\">Data leak?</a></p>",
      "rawMarkdown": "isakev there is a little mistake,sorry for that, we mean there are 600 validation samples that  are already in test data(not necessarily in public part) ,,, please check this [Data leak?](https://www.kaggle.com/sai11fkaneko/data-leak)",
      "votes": null
    },
    {
      "id": "891406",
      "postDate": "06/18/2020 06:52:09",
      "content": "<p><a href=\"/isakev\">@isakev</a> \n1) I am not 100% sure, but the test data we have access to is the public LB, right ?\n2) I assumed models were overfitting enough to remember the labels of the train samples. You may not get all guess correct but most of them will.</p>",
      "rawMarkdown": "isakev \n1) I am not 100% sure, but the test data we have access to is the public LB, right ?\n2) I assumed models were overfitting enough to remember the labels of the train samples. You may not get all guess correct but most of them will.",
      "votes": null
    },
    {
      "id": "891439",
      "postDate": "06/18/2020 07:23:38",
      "content": "<p><a href=\"/theoviel\">@theoviel</a> </p>\n\n<p>We have the full test data, there is no hidden private LB in this comp.</p>",
      "rawMarkdown": "theoviel \n\nWe have the full test data, there is no hidden private LB in this comp.",
      "votes": null
    },
    {
      "id": "891473",
      "postDate": "06/18/2020 07:40:57",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Thanks for pointing that out, I didn't follow carefully enough this competition. This changes what I said, the leak is not so bad then !</p>",
      "rawMarkdown": "philippsinger Thanks for pointing that out, I didn't follow carefully enough this competition. This changes what I said, the leak is not so bad then !",
      "votes": null
    },
    {
      "id": "894207",
      "postDate": "06/20/2020 08:25:56",
      "content": "<p>thank you very much </p>",
      "rawMarkdown": "thank you very much",
      "votes": null
    },
    {
      "id": "894242",
      "postDate": "06/20/2020 08:58:31",
      "content": "<p><a href=\"/theoviel\">@theoviel</a> : Any particular reason why hard coding the predictions to ground truth in validation set is not helping improve the LB score? (<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/154583\">See discussion</a>)</p>",
      "rawMarkdown": "theoviel : Any particular reason why hard coding the predictions to ground truth in validation set is not helping improve the LB score? ([See discussion](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/154583))",
      "votes": null
    },
    {
      "id": "894249",
      "postDate": "06/20/2020 09:06:43",
      "content": "<p>That's an intersting question.</p>\n\n<p>On the one hand, models are already good enough at capturing the information, and have correct predicitons for the \"leaky\" samples. </p>\n\n<p>On the other hand, some of the samples could have noisy labels, and hardcoding their predictions can hurt performances.</p>\n\n<p>But that's only hypothesis, so take that with a pinch of salt :) </p>",
      "rawMarkdown": "That's an intersting question.\n\nOn the one hand, models are already good enough at capturing the information, and have correct predicitons for the \"leaky\" samples. \n\nOn the other hand, some of the samples could have noisy labels, and hardcoding their predictions can hurt performances.\n\nBut that's only hypothesis, so take that with a pinch of salt :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 891000,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "06/17/2020 20:18:32",
      "content": "<p>Thanks for the info</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891174,
      "author_name": "veryrobustperson",
      "author_url": "",
      "post_date": "06/18/2020 01:23:20",
      "content": "<p>Very helpful. Thanks for the info.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891236,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "06/18/2020 03:26:01",
      "content": "<p>Thank you for sharing <a href=\"/theoviel\">@theoviel</a> . I must have missed something. \n1) How is it discovered that the 600 samples from validation.csv are in the public LB set and not partly in the private test set)?\n2) I thought that if a large model is trained on samples including these 600 then it would mostly classify correctly the copies of all those samples in the test set - so, no improvement in the PL score should be expected. Please correct me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 891355,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "06/18/2020 05:51:33",
          "content": "<p><a href=\"/isakev\">@isakev</a> there is a little mistake,sorry for that, we mean there are 600 validation samples that  are already in test data(not necessarily in public part) ,,, please check this <a href=\"https://www.kaggle.com/sai11fkaneko/data-leak\">Data leak?</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891406,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "06/18/2020 06:52:09",
          "content": "<p><a href=\"/isakev\">@isakev</a> \n1) I am not 100% sure, but the test data we have access to is the public LB, right ?\n2) I assumed models were overfitting enough to remember the labels of the train samples. You may not get all guess correct but most of them will.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891439,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "06/18/2020 07:23:38",
          "content": "<p><a href=\"/theoviel\">@theoviel</a> </p>\n\n<p>We have the full test data, there is no hidden private LB in this comp.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891473,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "06/18/2020 07:40:57",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Thanks for pointing that out, I didn't follow carefully enough this competition. This changes what I said, the leak is not so bad then !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 894207,
      "author_name": "sadikyetkin",
      "author_url": "",
      "post_date": "06/20/2020 08:25:56",
      "content": "<p>thank you very much </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 894242,
      "author_name": "nickil21",
      "author_url": "",
      "post_date": "06/20/2020 08:58:31",
      "content": "<p><a href=\"/theoviel\">@theoviel</a> : Any particular reason why hard coding the predictions to ground truth in validation set is not helping improve the LB score? (<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/154583\">See discussion</a>)</p>",
      "votes": null,
      "replies": [
        {
          "id": 894249,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "06/20/2020 09:06:43",
          "content": "<p>That's an intersting question.</p>\n\n<p>On the one hand, models are already good enough at capturing the information, and have correct predicitons for the \"leaky\" samples. </p>\n\n<p>On the other hand, some of the samples could have noisy labels, and hardcoding their predictions can hurt performances.</p>\n\n<p>But that's only hypothesis, so take that with a pinch of salt :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "890999": "We know that there are ~600 samples in the LB that are in the validation set.\n\nHowever, a lot of people are training on the validation set, which will results in over-estimated public LB scores. \n\nI wanted to check by how much approximately, and for a model that scores CV 0.945~, you can expect a boost varying from **0.001 to 0.002** from the leakage... Which is quite a lot considering how stacked the LB is.\n\nSo make sure you validate your models correctly ! \n\nEDIT : 600 samples in the complete LB, not public\n\nCode : \n```\nimport numpy as np\nfrom sklearn.metrics import *\nfrom tqdm.notebook import tqdm\n\nbase_scores = []\ndeltas = []\n\ntruth = np.concatenate((np.ones(7000), np.zeros(56000)))  # 63000 sample\nprint(f'Class 1 ratio : {truth.sum() / truth.size :.3f}') # 0.111\n\nfor i in tqdm(range(1000)):\n    pred = (truth / 2 + 0.25) + (np.random.random(size=truth.size) * 0.74 - 0.5) \n    base_score = roc_auc_score(truth, pred)\n\n    # Leakage : Setting 600 labels to their ground truth\n    pred[:70] = 0.99\n    pred[-530:] = 0.01\n\n    new_score = roc_auc_score(truth, pred)\n\n    base_scores.append(base_score)\n    deltas.append(new_score - base_score)\n\nprint(f'Average score : {np.mean(base_scores):.4f} - Average boost : {np.mean(deltas):.4f}')\n# Average score : 0.9474 - Average boost : 0.0010\n```",
    "891000": "Thanks for the info",
    "891174": "Very helpful. Thanks for the info.",
    "891236": "Thank you for sharing @theoviel . I must have missed something. \n1) How is it discovered that the 600 samples from validation.csv are in the public LB set and not partly in the private test set)?\n2) I thought that if a large model is trained on samples including these 600 then it would mostly classify correctly the copies of all those samples in the test set - so, no improvement in the PL score should be expected. Please correct me.",
    "891355": "isakev there is a little mistake,sorry for that, we mean there are 600 validation samples that  are already in test data(not necessarily in public part) ,,, please check this [Data leak?](https://www.kaggle.com/sai11fkaneko/data-leak)",
    "891406": "isakev \n1) I am not 100% sure, but the test data we have access to is the public LB, right ?\n2) I assumed models were overfitting enough to remember the labels of the train samples. You may not get all guess correct but most of them will.",
    "891439": "theoviel \n\nWe have the full test data, there is no hidden private LB in this comp.",
    "891473": "philippsinger Thanks for pointing that out, I didn't follow carefully enough this competition. This changes what I said, the leak is not so bad then !",
    "894207": "thank you very much",
    "894242": "theoviel : Any particular reason why hard coding the predictions to ground truth in validation set is not helping improve the LB score? ([See discussion](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/154583))",
    "894249": "That's an intersting question.\n\nOn the one hand, models are already good enough at capturing the information, and have correct predicitons for the \"leaky\" samples. \n\nOn the other hand, some of the samples could have noisy labels, and hardcoding their predictions can hurt performances.\n\nBut that's only hypothesis, so take that with a pinch of salt :)"
  },
  "source": "meta"
}