{
  "id": 79365,
  "title": "Be aware that some LB score is not reliable (trick to get 0.704)",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79365",
  "author_name": "",
  "post_date": "2019-02-03T04:13:29.872790600Z",
  "votes": 12,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I found a way to improve my LB score from 0.699 to 0.704, all you need is adding one line...</p>\n\n<blockquote>\n  <p>seed = random.randint(0, 99999)</p>\n</blockquote>\n\n<p>IMHO, trust in your CV score.\nHappy kaggling &amp; Chinese New Year.</p>",
  "messages": [
    {
      "id": "465421",
      "postDate": "02/03/2019 04:13:29",
      "content": "<p>I found a way to improve my LB score from 0.699 to 0.704, all you need is adding one line...</p>\n\n<blockquote>\n  <p>seed = random.randint(0, 99999)</p>\n</blockquote>\n\n<p>IMHO, trust in your CV score.\nHappy kaggling &amp; Chinese New Year.</p>",
      "rawMarkdown": "I found a way to improve my LB score from 0.699 to 0.704, all you need is adding one line...\n\n&gt; seed = random.randint(0, 99999)\n\nIMHO, trust in your CV score.\nHappy kaggling &amp; Chinese New Year.",
      "votes": null
    },
    {
      "id": "465425",
      "postDate": "02/03/2019 04:36:57",
      "content": "<p>You can also overfit to LB by moving the threshold around (since most popular way of setting threshold is by finding optimal threshold on train set and simply using the same threshold on test set, but that threshold will not necessarily be the best threshold for test set). </p>",
      "rawMarkdown": "You can also overfit to LB by moving the threshold around (since most popular way of setting threshold is by finding optimal threshold on train set and simply using the same threshold on test set, but that threshold will not necessarily be the best threshold for test set).",
      "votes": null
    },
    {
      "id": "465441",
      "postDate": "02/03/2019 05:31:45",
      "content": "<p>Is local CV reliable? My validation score improved from 0.682 to 0.687, but the public leader board still stuck at 0.700. Many people who have a high-rank report there local CV  report below 0.68. It would be shocking if they actually get a good position on the private leader board.</p>",
      "rawMarkdown": "Is local CV reliable? My validation score improved from 0.682 to 0.687, but the public leader board still stuck at 0.700. Many people who have a high-rank report there local CV  report below 0.68. It would be shocking if they actually get a good position on the private leader board.",
      "votes": null
    },
    {
      "id": "465523",
      "postDate": "02/03/2019 10:49:24",
      "content": "<p><img src=\"https://habrastorage.org/webt/vd/go/zs/vdgozs9pbyoroboyjfgup-rw5ju.jpeg\" alt=\"\"></p>",
      "rawMarkdown": "![](https://habrastorage.org/webt/vd/go/zs/vdgozs9pbyoroboyjfgup-rw5ju.jpeg)",
      "votes": null
    },
    {
      "id": "465592",
      "postDate": "02/03/2019 14:47:32",
      "content": "<p>Trust your local CV if you're using a validation set of significant size. The public LB score is based on only 56k samples, which is really tiny. </p>",
      "rawMarkdown": "Trust your local CV if you're using a validation set of significant size. The public LB score is based on only 56k samples, which is really tiny.",
      "votes": null
    },
    {
      "id": "465668",
      "postDate": "02/03/2019 18:33:25",
      "content": "<p>Go and add it to the \"Nerd Laughing Loud\" thread <a href=\"https://www.kaggle.com/general/76963\">https://www.kaggle.com/general/76963</a></p>",
      "rawMarkdown": "Go and add it to the \"Nerd Laughing Loud\" thread https://www.kaggle.com/general/76963",
      "votes": null
    },
    {
      "id": "465718",
      "postDate": "02/03/2019 20:54:21",
      "content": "<p>I've been reading a lot solutions to the past competitions. One thing I constantly hear about is to set the threshold so that the train and test set have the same distribution. I am little confused about the definition of \"same distribution\". Does it mean that we want the test set has the same percentage of positive labels as in train set?</p>",
      "rawMarkdown": "I've been reading a lot solutions to the past competitions. One thing I constantly hear about is to set the threshold so that the train and test set have the same distribution. I am little confused about the definition of \"same distribution\". Does it mean that we want the test set has the same percentage of positive labels as in train set?",
      "votes": null
    },
    {
      "id": "465776",
      "postDate": "02/04/2019 01:32:34",
      "content": "<p>The size of LB is so small, the small difference between LB scores could be caused by the small size.  Higher cv models should get higher score when the size of testing set become much larger and the distribution is closer to the distribution of training set. Using k fold cv and taking an average prediction of k models would increase your models' ability.  I think it's possible to get LB around 0.7 by the model with local cv around 0.68. But higher LB may use some tricks or different methods from the common public kernel. IMO, local cv is more reliable than LB.</p>",
      "rawMarkdown": "The size of LB is so small, the small difference between LB scores could be caused by the small size.  Higher cv models should get higher score when the size of testing set become much larger and the distribution is closer to the distribution of training set. Using k fold cv and taking an average prediction of k models would increase your models' ability.  I think it's possible to get LB around 0.7 by the model with local cv around 0.68. But higher LB may use some tricks or different methods from the common public kernel. IMO, local cv is more reliable than LB.",
      "votes": null
    },
    {
      "id": "465779",
      "postDate": "02/04/2019 01:40:39",
      "content": "<p>Happy Chinese New Year!!  It seems that you get a magic seed, is your result reproducible? Any difference for Local CV?  Thanks!</p>",
      "rawMarkdown": "Happy Chinese New Year!!  It seems that you get a magic seed, is your result reproducible? Any difference for Local CV?  Thanks!",
      "votes": null
    },
    {
      "id": "465786",
      "postDate": "02/04/2019 02:37:50",
      "content": "<p>Honestly, it's completely unknown :^). We don't really know the details of Quora's new test set. Assuming that they set this competition up properly, their test set distribution should reflect the training data they gave us (otherwise it's impossible to guarantee reasonable results). However, theoretically they could provide some completely arbitrary test dataset :)</p>",
      "rawMarkdown": "Honestly, it's completely unknown :^). We don't really know the details of Quora's new test set. Assuming that they set this competition up properly, their test set distribution should reflect the training data they gave us (otherwise it's impossible to guarantee reasonable results). However, theoretically they could provide some completely arbitrary test dataset :)",
      "votes": null
    },
    {
      "id": "465817",
      "postDate": "02/04/2019 04:28:30",
      "content": "<p>Local CV is much better than LB, though even local CV isn't completely accurate at predicting performance on the final test set. The blogpost linked below talks about the author's entries to a previous kaggle competition and how even though local CV score improved, private score was not improving with the CV score. Public score was all over the place.</p>\n\n<p>I don't think the highest scoring people will necessarily submit their highest LB kernels. They should have other options that have better local CV scores.</p>\n\n<p><a href=\"http://gregpark.io/blog/Kaggle-Psychopathy-Postmortem/\">http://gregpark.io/blog/Kaggle-Psychopathy-Postmortem/</a></p>",
      "rawMarkdown": "Local CV is much better than LB, though even local CV isn't completely accurate at predicting performance on the final test set. The blogpost linked below talks about the author's entries to a previous kaggle competition and how even though local CV score improved, private score was not improving with the CV score. Public score was all over the place.\n\nI don't think the highest scoring people will necessarily submit their highest LB kernels. They should have other options that have better local CV scores.\n\n\nhttp://gregpark.io/blog/Kaggle-Psychopathy-Postmortem/",
      "votes": null
    },
    {
      "id": "465818",
      "postDate": "02/04/2019 04:36:57",
      "content": "<p>I believe that is what they mean. The recommendation is probably running off the assumption that the train and test sets should have approximately the same proportions of each class. It isn't something I've tried.</p>",
      "rawMarkdown": "I believe that is what they mean. The recommendation is probably running off the assumption that the train and test sets should have approximately the same proportions of each class. It isn't something I've tried.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 465425,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "02/03/2019 04:36:57",
      "content": "<p>You can also overfit to LB by moving the threshold around (since most popular way of setting threshold is by finding optimal threshold on train set and simply using the same threshold on test set, but that threshold will not necessarily be the best threshold for test set). </p>",
      "votes": null,
      "replies": [
        {
          "id": 465718,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "02/03/2019 20:54:21",
          "content": "<p>I've been reading a lot solutions to the past competitions. One thing I constantly hear about is to set the threshold so that the train and test set have the same distribution. I am little confused about the definition of \"same distribution\". Does it mean that we want the test set has the same percentage of positive labels as in train set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465818,
          "author_name": "julius6",
          "author_url": "",
          "post_date": "02/04/2019 04:36:57",
          "content": "<p>I believe that is what they mean. The recommendation is probably running off the assumption that the train and test sets should have approximately the same proportions of each class. It isn't something I've tried.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465441,
      "author_name": "wenrui29",
      "author_url": "",
      "post_date": "02/03/2019 05:31:45",
      "content": "<p>Is local CV reliable? My validation score improved from 0.682 to 0.687, but the public leader board still stuck at 0.700. Many people who have a high-rank report there local CV  report below 0.68. It would be shocking if they actually get a good position on the private leader board.</p>",
      "votes": null,
      "replies": [
        {
          "id": 465592,
          "author_name": "gamerx",
          "author_url": "",
          "post_date": "02/03/2019 14:47:32",
          "content": "<p>Trust your local CV if you're using a validation set of significant size. The public LB score is based on only 56k samples, which is really tiny. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465776,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "02/04/2019 01:32:34",
          "content": "<p>The size of LB is so small, the small difference between LB scores could be caused by the small size.  Higher cv models should get higher score when the size of testing set become much larger and the distribution is closer to the distribution of training set. Using k fold cv and taking an average prediction of k models would increase your models' ability.  I think it's possible to get LB around 0.7 by the model with local cv around 0.68. But higher LB may use some tricks or different methods from the common public kernel. IMO, local cv is more reliable than LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465786,
          "author_name": "msf908",
          "author_url": "",
          "post_date": "02/04/2019 02:37:50",
          "content": "<p>Honestly, it's completely unknown :^). We don't really know the details of Quora's new test set. Assuming that they set this competition up properly, their test set distribution should reflect the training data they gave us (otherwise it's impossible to guarantee reasonable results). However, theoretically they could provide some completely arbitrary test dataset :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465817,
          "author_name": "julius6",
          "author_url": "",
          "post_date": "02/04/2019 04:28:30",
          "content": "<p>Local CV is much better than LB, though even local CV isn't completely accurate at predicting performance on the final test set. The blogpost linked below talks about the author's entries to a previous kaggle competition and how even though local CV score improved, private score was not improving with the CV score. Public score was all over the place.</p>\n\n<p>I don't think the highest scoring people will necessarily submit their highest LB kernels. They should have other options that have better local CV scores.</p>\n\n<p><a href=\"http://gregpark.io/blog/Kaggle-Psychopathy-Postmortem/\">http://gregpark.io/blog/Kaggle-Psychopathy-Postmortem/</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465523,
      "author_name": "",
      "author_url": "",
      "post_date": "02/03/2019 10:49:24",
      "content": "<p><img src=\"https://habrastorage.org/webt/vd/go/zs/vdgozs9pbyoroboyjfgup-rw5ju.jpeg\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 465668,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "02/03/2019 18:33:25",
          "content": "<p>Go and add it to the \"Nerd Laughing Loud\" thread <a href=\"https://www.kaggle.com/general/76963\">https://www.kaggle.com/general/76963</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 465779,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "02/04/2019 01:40:39",
      "content": "<p>Happy Chinese New Year!!  It seems that you get a magic seed, is your result reproducible? Any difference for Local CV?  Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "465421": "I found a way to improve my LB score from 0.699 to 0.704, all you need is adding one line...\n\n&gt; seed = random.randint(0, 99999)\n\nIMHO, trust in your CV score.\nHappy kaggling &amp; Chinese New Year.",
    "465425": "You can also overfit to LB by moving the threshold around (since most popular way of setting threshold is by finding optimal threshold on train set and simply using the same threshold on test set, but that threshold will not necessarily be the best threshold for test set).",
    "465441": "Is local CV reliable? My validation score improved from 0.682 to 0.687, but the public leader board still stuck at 0.700. Many people who have a high-rank report there local CV  report below 0.68. It would be shocking if they actually get a good position on the private leader board.",
    "465523": "![](https://habrastorage.org/webt/vd/go/zs/vdgozs9pbyoroboyjfgup-rw5ju.jpeg)",
    "465592": "Trust your local CV if you're using a validation set of significant size. The public LB score is based on only 56k samples, which is really tiny.",
    "465668": "Go and add it to the \"Nerd Laughing Loud\" thread https://www.kaggle.com/general/76963",
    "465718": "I've been reading a lot solutions to the past competitions. One thing I constantly hear about is to set the threshold so that the train and test set have the same distribution. I am little confused about the definition of \"same distribution\". Does it mean that we want the test set has the same percentage of positive labels as in train set?",
    "465776": "The size of LB is so small, the small difference between LB scores could be caused by the small size.  Higher cv models should get higher score when the size of testing set become much larger and the distribution is closer to the distribution of training set. Using k fold cv and taking an average prediction of k models would increase your models' ability.  I think it's possible to get LB around 0.7 by the model with local cv around 0.68. But higher LB may use some tricks or different methods from the common public kernel. IMO, local cv is more reliable than LB.",
    "465779": "Happy Chinese New Year!!  It seems that you get a magic seed, is your result reproducible? Any difference for Local CV?  Thanks!",
    "465786": "Honestly, it's completely unknown :^). We don't really know the details of Quora's new test set. Assuming that they set this competition up properly, their test set distribution should reflect the training data they gave us (otherwise it's impossible to guarantee reasonable results). However, theoretically they could provide some completely arbitrary test dataset :)",
    "465817": "Local CV is much better than LB, though even local CV isn't completely accurate at predicting performance on the final test set. The blogpost linked below talks about the author's entries to a previous kaggle competition and how even though local CV score improved, private score was not improving with the CV score. Public score was all over the place.\n\nI don't think the highest scoring people will necessarily submit their highest LB kernels. They should have other options that have better local CV scores.\n\n\nhttp://gregpark.io/blog/Kaggle-Psychopathy-Postmortem/",
    "465818": "I believe that is what they mean. The recommendation is probably running off the assumption that the train and test sets should have approximately the same proportions of each class. It isn't something I've tried."
  },
  "source": "meta"
}