{
  "id": 71241,
  "title": "Thresholds and Validation F1 Score",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71241",
  "author_name": "",
  "post_date": "2018-11-11T19:50:57.934310800Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm seeing a lot of kernels performing a final threshold analysis to pick the threshold for the highest validation f1-score. It looks a bit like this:</p>\n\n<pre><code>pred_val_y = 0.6 * pred_val_lstm_y + 0.4 * pred_val_cnn_y  # two random numbers :)\npred_test_y = 0.6 * pred_test_lstm_y + 0.4 * pred_test_cnn_y \n\nthresholds = []\nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    res = metrics.f1_score(val_y, (pred_val_y &amp;gt; thresh).astype(int))\n    thresholds.append([thresh, res])\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, res))\n\nthresholds.sort(key=lambda x: x[1], reverse=True)\nbest_thresh = thresholds[0][0]\nprint(\"Best threshold: \", best_thresh)\n</code></pre>\n\n<p>(from <a href=\"https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d\">https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d</a>)</p>\n\n<p>I'm a little confused why this is used so much in kernels with NN's with a single sigmoid output. I am having difficulty understanding why we do this instead of just fixing 0.5 as the threshold. Is this due to the data imbalance?</p>\n\n<p>Andrew</p>",
  "messages": [
    {
      "id": "419364",
      "postDate": "11/11/2018 19:50:57",
      "content": "<p>I'm seeing a lot of kernels performing a final threshold analysis to pick the threshold for the highest validation f1-score. It looks a bit like this:</p>\n\n<pre><code>pred_val_y = 0.6 * pred_val_lstm_y + 0.4 * pred_val_cnn_y  # two random numbers :)\npred_test_y = 0.6 * pred_test_lstm_y + 0.4 * pred_test_cnn_y \n\nthresholds = []\nfor thresh in np.arange(0.1, 0.501, 0.01):\n    thresh = np.round(thresh, 2)\n    res = metrics.f1_score(val_y, (pred_val_y &amp;gt; thresh).astype(int))\n    thresholds.append([thresh, res])\n    print(\"F1 score at threshold {0} is {1}\".format(thresh, res))\n\nthresholds.sort(key=lambda x: x[1], reverse=True)\nbest_thresh = thresholds[0][0]\nprint(\"Best threshold: \", best_thresh)\n</code></pre>\n\n<p>(from <a href=\"https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d\">https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d</a>)</p>\n\n<p>I'm a little confused why this is used so much in kernels with NN's with a single sigmoid output. I am having difficulty understanding why we do this instead of just fixing 0.5 as the threshold. Is this due to the data imbalance?</p>\n\n<p>Andrew</p>",
      "rawMarkdown": "I'm seeing a lot of kernels performing a final threshold analysis to pick the threshold for the highest validation f1-score. It looks a bit like this:\n\n    pred_val_y = 0.6 * pred_val_lstm_y + 0.4 * pred_val_cnn_y  # two random numbers :)\n    pred_test_y = 0.6 * pred_test_lstm_y + 0.4 * pred_test_cnn_y \n    \n    thresholds = []\n    for thresh in np.arange(0.1, 0.501, 0.01):\n        thresh = np.round(thresh, 2)\n        res = metrics.f1_score(val_y, (pred_val_y &gt; thresh).astype(int))\n        thresholds.append([thresh, res])\n        print(\"F1 score at threshold {0} is {1}\".format(thresh, res))\n        \n    thresholds.sort(key=lambda x: x[1], reverse=True)\n    best_thresh = thresholds[0][0]\n    print(\"Best threshold: \", best_thresh)\n\n(from https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d)\n\nI'm a little confused why this is used so much in kernels with NN's with a single sigmoid output. I am having difficulty understanding why we do this instead of just fixing 0.5 as the threshold. Is this due to the data imbalance?\n\nAndrew",
      "votes": null
    },
    {
      "id": "419366",
      "postDate": "11/11/2018 20:09:50",
      "content": "<p>since I wrote this post, I discovered this paper: <a href=\"https://arxiv.org/pdf/1402.1892.pdf\">https://arxiv.org/pdf/1402.1892.pdf</a></p>",
      "rawMarkdown": "since I wrote this post, I discovered this paper: https://arxiv.org/pdf/1402.1892.pdf",
      "votes": null
    },
    {
      "id": "419513",
      "postDate": "11/12/2018 05:10:16",
      "content": "<p>I think data imbalance is one reason, and moreover 0.5 is a standard threshold right. It all depends on the problem, in some problem reducing False Negative is the consideration, in some others reducing False Positive is the preference. So, I think different thresholds might have different results and impact. </p>",
      "rawMarkdown": "I think data imbalance is one reason, and moreover 0.5 is a standard threshold right. It all depends on the problem, in some problem reducing False Negative is the consideration, in some others reducing False Positive is the preference. So, I think different thresholds might have different results and impact.",
      "votes": null
    },
    {
      "id": "420312",
      "postDate": "11/13/2018 12:47:50",
      "content": "<p>There is a better way:  <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71184#latest-419289\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71184#latest-419289</a></p>",
      "rawMarkdown": "There is a better way:  https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71184#latest-419289",
      "votes": null
    },
    {
      "id": "585844",
      "postDate": "07/28/2019 07:31:17",
      "content": "<p>Any idea how to do F1 score threshold for multiclass classification problem?</p>",
      "rawMarkdown": "Any idea how to do F1 score threshold for multiclass classification problem?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419366,
      "author_name": "",
      "author_url": "",
      "post_date": "11/11/2018 20:09:50",
      "content": "<p>since I wrote this post, I discovered this paper: <a href=\"https://arxiv.org/pdf/1402.1892.pdf\">https://arxiv.org/pdf/1402.1892.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419513,
      "author_name": "s4sarath",
      "author_url": "",
      "post_date": "11/12/2018 05:10:16",
      "content": "<p>I think data imbalance is one reason, and moreover 0.5 is a standard threshold right. It all depends on the problem, in some problem reducing False Negative is the consideration, in some others reducing False Positive is the preference. So, I think different thresholds might have different results and impact. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420312,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/13/2018 12:47:50",
      "content": "<p>There is a better way:  <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71184#latest-419289\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71184#latest-419289</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 585844,
      "author_name": "shravankoninti",
      "author_url": "",
      "post_date": "07/28/2019 07:31:17",
      "content": "<p>Any idea how to do F1 score threshold for multiclass classification problem?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "419364": "I'm seeing a lot of kernels performing a final threshold analysis to pick the threshold for the highest validation f1-score. It looks a bit like this:\n\n    pred_val_y = 0.6 * pred_val_lstm_y + 0.4 * pred_val_cnn_y  # two random numbers :)\n    pred_test_y = 0.6 * pred_test_lstm_y + 0.4 * pred_test_cnn_y \n    \n    thresholds = []\n    for thresh in np.arange(0.1, 0.501, 0.01):\n        thresh = np.round(thresh, 2)\n        res = metrics.f1_score(val_y, (pred_val_y &gt; thresh).astype(int))\n        thresholds.append([thresh, res])\n        print(\"F1 score at threshold {0} is {1}\".format(thresh, res))\n        \n    thresholds.sort(key=lambda x: x[1], reverse=True)\n    best_thresh = thresholds[0][0]\n    print(\"Best threshold: \", best_thresh)\n\n(from https://www.kaggle.com/shujian/blend-of-lstm-and-cnn-with-4-embeddings-1200d)\n\nI'm a little confused why this is used so much in kernels with NN's with a single sigmoid output. I am having difficulty understanding why we do this instead of just fixing 0.5 as the threshold. Is this due to the data imbalance?\n\nAndrew",
    "419366": "since I wrote this post, I discovered this paper: https://arxiv.org/pdf/1402.1892.pdf",
    "419513": "I think data imbalance is one reason, and moreover 0.5 is a standard threshold right. It all depends on the problem, in some problem reducing False Negative is the consideration, in some others reducing False Positive is the preference. So, I think different thresholds might have different results and impact.",
    "420312": "There is a better way:  https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71184#latest-419289",
    "585844": "Any idea how to do F1 score threshold for multiclass classification problem?"
  },
  "source": "meta"
}