{
  "id": 77076,
  "title": "CV before and after threshold ",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77076",
  "author_name": "",
  "post_date": "2019-01-09T09:08:45.984171100Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am curious to know what's your cv score at 0.5 threshold and optimal threshold.</p>\n\n<p>Let me start with mine : \n1) cv score of 0.6742 and after finding the optimal threshold my cv score is 0.6910 at threshold 0.38 (which is a huge jump).</p>\n\n<p>2) In my another kernel the cv score is 0.6764 and after finding optimal threshold my score jumped to 0.6918 at 0.39 threshold</p>",
  "messages": [
    {
      "id": "452852",
      "postDate": "01/09/2019 09:08:45",
      "content": "<p>I am curious to know what's your cv score at 0.5 threshold and optimal threshold.</p>\n\n<p>Let me start with mine : \n1) cv score of 0.6742 and after finding the optimal threshold my cv score is 0.6910 at threshold 0.38 (which is a huge jump).</p>\n\n<p>2) In my another kernel the cv score is 0.6764 and after finding optimal threshold my score jumped to 0.6918 at 0.39 threshold</p>",
      "rawMarkdown": "I am curious to know what's your cv score at 0.5 threshold and optimal threshold.\n\nLet me start with mine : \n1) cv score of 0.6742 and after finding the optimal threshold my cv score is 0.6910 at threshold 0.38 (which is a huge jump).\n\n2) In my another kernel the cv score is 0.6764 and after finding optimal threshold my score jumped to 0.6918 at 0.39 threshold",
      "votes": null
    },
    {
      "id": "453789",
      "postDate": "01/10/2019 19:35:21",
      "content": "<p>I think there is no reason for the model to have a threshold at 0.5, because the negative data is much more than the positive ones. Many high score public kernel set their threshold at 0.33 or 0.34, rather than use the threshold search. \nBut actually when I test the influence of the threshold, I find the wide-used RNN models are quite robust for this hyper parameter. I set the threshold 0.1 less or more than the best threshold, and the f1 score just become lower no more than 0.003, which is kind of acceptable regarding the large difference between CV and LB may because of  overfitting the threshold. </p>",
      "rawMarkdown": "I think there is no reason for the model to have a threshold at 0.5, because the negative data is much more than the positive ones. Many high score public kernel set their threshold at 0.33 or 0.34, rather than use the threshold search. \nBut actually when I test the influence of the threshold, I find the wide-used RNN models are quite robust for this hyper parameter. I set the threshold 0.1 less or more than the best threshold, and the f1 score just become lower no more than 0.003, which is kind of acceptable regarding the large difference between CV and LB may because of  overfitting the threshold.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 453789,
      "author_name": "peining",
      "author_url": "",
      "post_date": "01/10/2019 19:35:21",
      "content": "<p>I think there is no reason for the model to have a threshold at 0.5, because the negative data is much more than the positive ones. Many high score public kernel set their threshold at 0.33 or 0.34, rather than use the threshold search. \nBut actually when I test the influence of the threshold, I find the wide-used RNN models are quite robust for this hyper parameter. I set the threshold 0.1 less or more than the best threshold, and the f1 score just become lower no more than 0.003, which is kind of acceptable regarding the large difference between CV and LB may because of  overfitting the threshold. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "452852": "I am curious to know what's your cv score at 0.5 threshold and optimal threshold.\n\nLet me start with mine : \n1) cv score of 0.6742 and after finding the optimal threshold my cv score is 0.6910 at threshold 0.38 (which is a huge jump).\n\n2) In my another kernel the cv score is 0.6764 and after finding optimal threshold my score jumped to 0.6918 at 0.39 threshold",
    "453789": "I think there is no reason for the model to have a threshold at 0.5, because the negative data is much more than the positive ones. Many high score public kernel set their threshold at 0.33 or 0.34, rather than use the threshold search. \nBut actually when I test the influence of the threshold, I find the wide-used RNN models are quite robust for this hyper parameter. I set the threshold 0.1 less or more than the best threshold, and the f1 score just become lower no more than 0.003, which is kind of acceptable regarding the large difference between CV and LB may because of  overfitting the threshold."
  },
  "source": "meta"
}