{
  "id": 79556,
  "title": "Fitting tokenizers with validation data might leak information to model",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79556",
  "author_name": "bilal2vec",
  "post_date": "2019-02-05T13:42:23.478000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I've noticed that most of the public kernels fit the tokenizer on the entire train.csv file, this may be artificially increasing the local cv score by lowering the number of oov words in each fold's validation data. </p>\n\n<p>This could be a problem since the tokenizer is fit on data that will be used for validating the model after training each fold and as a result will not cover as much of the test set's vocabulary while covering more of the validation set and increasing the local cv score. This might also increase the number of oov words being passed to the model when the kernels are run with the new stage 2 test set and end up lowering the score.</p>\n\n<p>As pointed out in <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79459\">this discussion post</a>, fitting the tokenizer on the test set as well as the full train set won't fix the problem and might make the problem worse. One possible way to avoid leaking information to the model could be to fit a new tokenizer to each fold's training data and load in a new set of embeddings at each fold. This should stop the leakage but will increase the run time by a lot.</p>",
  "messages": [
    {
      "id": 466493,
      "postDate": "2019-02-05T13:42:23.480Z",
      "content": "<p>Hi,</p>\n\n<p>I've noticed that most of the public kernels fit the tokenizer on the entire train.csv file, this may be artificially increasing the local cv score by lowering the number of oov words in each fold's validation data. </p>\n\n<p>This could be a problem since the tokenizer is fit on data that will be used for validating the model after training each fold and as a result will not cover as much of the test set's vocabulary while covering more of the validation set and increasing the local cv score. This might also increase the number of oov words being passed to the model when the kernels are run with the new stage 2 test set and end up lowering the score.</p>\n\n<p>As pointed out in <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79459\">this discussion post</a>, fitting the tokenizer on the test set as well as the full train set won't fix the problem and might make the problem worse. One possible way to avoid leaking information to the model could be to fit a new tokenizer to each fold's training data and load in a new set of embeddings at each fold. This should stop the leakage but will increase the run time by a lot.</p>",
      "rawMarkdown": "Hi,\n\nI've noticed that most of the public kernels fit the tokenizer on the entire train.csv file, this may be artificially increasing the local cv score by lowering the number of oov words in each fold's validation data. \n\nThis could be a problem since the tokenizer is fit on data that will be used for validating the model after training each fold and as a result will not cover as much of the test set's vocabulary while covering more of the validation set and increasing the local cv score. This might also increase the number of oov words being passed to the model when the kernels are run with the new stage 2 test set and end up lowering the score.\n\nAs pointed out in [this discussion post][1], fitting the tokenizer on the test set as well as the full train set won't fix the problem and might make the problem worse. One possible way to avoid leaking information to the model could be to fit a new tokenizer to each fold's training data and load in a new set of embeddings at each fold. This should stop the leakage but will increase the run time by a lot.\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79459",
      "votes": 3
    },
    {
      "id": 466532,
      "postDate": "2019-02-05T15:12:53.743Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 466532,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-05T15:12:53.743000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "466493": "Hi,\n\nI've noticed that most of the public kernels fit the tokenizer on the entire train.csv file, this may be artificially increasing the local cv score by lowering the number of oov words in each fold's validation data. \n\nThis could be a problem since the tokenizer is fit on data that will be used for validating the model after training each fold and as a result will not cover as much of the test set's vocabulary while covering more of the validation set and increasing the local cv score. This might also increase the number of oov words being passed to the model when the kernels are run with the new stage 2 test set and end up lowering the score.\n\nAs pointed out in [this discussion post][1], fitting the tokenizer on the test set as well as the full train set won't fix the problem and might make the problem worse. One possible way to avoid leaking information to the model could be to fit a new tokenizer to each fold's training data and load in a new set of embeddings at each fold. This should stop the leakage but will increase the run time by a lot.\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79459",
    "466532": ""
  }
}