{
  "id": 75735,
  "title": "Faster approach for finding the optimal threshold ",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75735",
  "author_name": "",
  "post_date": "2018-12-25T17:40:18.030062900Z",
  "votes": 86,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hey. </p>\n\n<p>I've noticed that a lot of kernels contain <code>for</code> loop in optimal threshold search. <code>for</code> loops are inefficient and you may replace them with faster sklearn implementation. </p>\n\n<p>Kernel's approach:</p>\n\n<pre><code>def threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in [i * 0.01 for i in range(100)]:\n        score = f1_score(y_true=y_true, y_pred=y_proba &amp;gt; threshold)\n        if score &amp;gt; best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'f1': best_score}\n    return search_result\n</code></pre>\n\n<p>Better approach: </p>\n\n<pre><code>from sklearn.metrics import roc_curve, precision_recall_curve\ndef threshold_search(y_true, y_proba, plot=False):\n    precision, recall, thresholds = precision_recall_curve(y_true, y_proba)\n    thresholds = np.append(thresholds, 1.001) \n    F = 2 / (1/precision + 1/recall)\n    best_score = np.max(F)\n    best_th = thresholds[np.argmax(F)]\n    if plot:\n        plt.plot(thresholds, F, '-b')\n        plt.plot([best_th], [best_score], '*r')\n        plt.show()\n    search_result = {'threshold': best_th , 'f1': best_score}\n    return search_result \n</code></pre>\n\n<p>Time comparing: \nNumber of probs = 200k, first script - 4.01s, second script - 70ms. </p>\n\n<p>Happy kaggling! </p>\n\n<p>UPD: A better approach for finding the optimal threshold \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958</a></p>",
  "messages": [
    {
      "id": "445122",
      "postDate": "12/25/2018 17:40:18",
      "content": "<p>Hey. </p>\n\n<p>I've noticed that a lot of kernels contain <code>for</code> loop in optimal threshold search. <code>for</code> loops are inefficient and you may replace them with faster sklearn implementation. </p>\n\n<p>Kernel's approach:</p>\n\n<pre><code>def threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in [i * 0.01 for i in range(100)]:\n        score = f1_score(y_true=y_true, y_pred=y_proba &amp;gt; threshold)\n        if score &amp;gt; best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'f1': best_score}\n    return search_result\n</code></pre>\n\n<p>Better approach: </p>\n\n<pre><code>from sklearn.metrics import roc_curve, precision_recall_curve\ndef threshold_search(y_true, y_proba, plot=False):\n    precision, recall, thresholds = precision_recall_curve(y_true, y_proba)\n    thresholds = np.append(thresholds, 1.001) \n    F = 2 / (1/precision + 1/recall)\n    best_score = np.max(F)\n    best_th = thresholds[np.argmax(F)]\n    if plot:\n        plt.plot(thresholds, F, '-b')\n        plt.plot([best_th], [best_score], '*r')\n        plt.show()\n    search_result = {'threshold': best_th , 'f1': best_score}\n    return search_result \n</code></pre>\n\n<p>Time comparing: \nNumber of probs = 200k, first script - 4.01s, second script - 70ms. </p>\n\n<p>Happy kaggling! </p>\n\n<p>UPD: A better approach for finding the optimal threshold \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958</a></p>",
      "rawMarkdown": "Hey. \n\nI've noticed that a lot of kernels contain `for` loop in optimal threshold search. `for` loops are inefficient and you may replace them with faster sklearn implementation. \n\nKernel's approach:\n\n    def threshold_search(y_true, y_proba):\n        best_threshold = 0\n        best_score = 0\n        for threshold in [i * 0.01 for i in range(100)]:\n            score = f1_score(y_true=y_true, y_pred=y_proba &gt; threshold)\n            if score &gt; best_score:\n                best_threshold = threshold\n                best_score = score\n        search_result = {'threshold': best_threshold, 'f1': best_score}\n        return search_result\n\nBetter approach: \n\n    from sklearn.metrics import roc_curve, precision_recall_curve\n    def threshold_search(y_true, y_proba, plot=False):\n        precision, recall, thresholds = precision_recall_curve(y_true, y_proba)\n        thresholds = np.append(thresholds, 1.001) \n        F = 2 / (1/precision + 1/recall)\n        best_score = np.max(F)\n        best_th = thresholds[np.argmax(F)]\n        if plot:\n            plt.plot(thresholds, F, '-b')\n            plt.plot([best_th], [best_score], '*r')\n            plt.show()\n        search_result = {'threshold': best_th , 'f1': best_score}\n        return search_result \n\nTime comparing: \nNumber of probs = 200k, first script - 4.01s, second script - 70ms. \n\nHappy kaggling! \n\nUPD: A better approach for finding the optimal threshold \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958",
      "votes": null
    },
    {
      "id": "445128",
      "postDate": "12/25/2018 18:02:54",
      "content": "<p>Should optimize my runtime even more. Thanks for this!</p>",
      "rawMarkdown": "Should optimize my runtime even more. Thanks for this!",
      "votes": null
    },
    {
      "id": "445238",
      "postDate": "12/26/2018 03:21:42",
      "content": "<p>Thanks for pointing out this point where I usually overlooked.</p>",
      "rawMarkdown": "Thanks for pointing out this point where I usually overlooked.",
      "votes": null
    },
    {
      "id": "445645",
      "postDate": "12/26/2018 21:29:14",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "445748",
      "postDate": "12/27/2018 02:47:47",
      "content": "<p>Nice . Thank you.</p>",
      "rawMarkdown": "Nice . Thank you.",
      "votes": null
    },
    {
      "id": "445850",
      "postDate": "12/27/2018 05:53:57",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "445940",
      "postDate": "12/27/2018 08:47:47",
      "content": "<p>Thank you. Quite valuable. </p>",
      "rawMarkdown": "Thank you. Quite valuable.",
      "votes": null
    },
    {
      "id": "445963",
      "postDate": "12/27/2018 09:23:43",
      "content": "<p>Thanks for bringing this up. Very useful :)</p>",
      "rawMarkdown": "Thanks for bringing this up. Very useful :)",
      "votes": null
    },
    {
      "id": "449024",
      "postDate": "01/02/2019 14:43:35",
      "content": "<p>UPD: A better approach for finding the optimal threshold \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958</a></p>",
      "rawMarkdown": "UPD: A better approach for finding the optimal threshold \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958",
      "votes": null
    },
    {
      "id": "1581177",
      "postDate": "11/13/2021 13:30:13",
      "content": "<p>Thanks for sharing ! Very useful !</p>",
      "rawMarkdown": "Thanks for sharing ! Very useful !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1581177,
      "author_name": "boullitmohamedfaycal",
      "author_url": "",
      "post_date": "11/13/2021 13:30:13",
      "content": "<p>Thanks for sharing ! Very useful !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445128,
      "author_name": "shaz13",
      "author_url": "",
      "post_date": "12/25/2018 18:02:54",
      "content": "<p>Should optimize my runtime even more. Thanks for this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445238,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "12/26/2018 03:21:42",
      "content": "<p>Thanks for pointing out this point where I usually overlooked.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445645,
      "author_name": "spirosrap",
      "author_url": "",
      "post_date": "12/26/2018 21:29:14",
      "content": "<p>Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445748,
      "author_name": "s4sarath",
      "author_url": "",
      "post_date": "12/27/2018 02:47:47",
      "content": "<p>Nice . Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445850,
      "author_name": "lingzz",
      "author_url": "",
      "post_date": "12/27/2018 05:53:57",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445940,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "12/27/2018 08:47:47",
      "content": "<p>Thank you. Quite valuable. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445963,
      "author_name": "naveenkumar1",
      "author_url": "",
      "post_date": "12/27/2018 09:23:43",
      "content": "<p>Thanks for bringing this up. Very useful :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449024,
      "author_name": "",
      "author_url": "",
      "post_date": "01/02/2019 14:43:35",
      "content": "<p>UPD: A better approach for finding the optimal threshold \n<a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "445122": "Hey. \n\nI've noticed that a lot of kernels contain `for` loop in optimal threshold search. `for` loops are inefficient and you may replace them with faster sklearn implementation. \n\nKernel's approach:\n\n    def threshold_search(y_true, y_proba):\n        best_threshold = 0\n        best_score = 0\n        for threshold in [i * 0.01 for i in range(100)]:\n            score = f1_score(y_true=y_true, y_pred=y_proba &gt; threshold)\n            if score &gt; best_score:\n                best_threshold = threshold\n                best_score = score\n        search_result = {'threshold': best_threshold, 'f1': best_score}\n        return search_result\n\nBetter approach: \n\n    from sklearn.metrics import roc_curve, precision_recall_curve\n    def threshold_search(y_true, y_proba, plot=False):\n        precision, recall, thresholds = precision_recall_curve(y_true, y_proba)\n        thresholds = np.append(thresholds, 1.001) \n        F = 2 / (1/precision + 1/recall)\n        best_score = np.max(F)\n        best_th = thresholds[np.argmax(F)]\n        if plot:\n            plt.plot(thresholds, F, '-b')\n            plt.plot([best_th], [best_score], '*r')\n            plt.show()\n        search_result = {'threshold': best_th , 'f1': best_score}\n        return search_result \n\nTime comparing: \nNumber of probs = 200k, first script - 4.01s, second script - 70ms. \n\nHappy kaggling! \n\nUPD: A better approach for finding the optimal threshold \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958",
    "445128": "Should optimize my runtime even more. Thanks for this!",
    "445238": "Thanks for pointing out this point where I usually overlooked.",
    "445645": "Thank you!",
    "445748": "Nice . Thank you.",
    "445850": "Thanks for sharing!",
    "445940": "Thank you. Quite valuable.",
    "445963": "Thanks for bringing this up. Very useful :)",
    "449024": "UPD: A better approach for finding the optimal threshold \nhttps://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76391#448958",
    "1581177": "Thanks for sharing ! Very useful !"
  },
  "source": "meta"
}