{
  "id": 79179,
  "title": "Fast threshold calculation",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79179",
  "author_name": "karthik kolli",
  "post_date": "2019-02-01T01:06:59.708000",
  "votes": -1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>This is a fast implementation to calculate threshold. Instead of calling f1_score manually, this does similar job. </p>\n\n<pre><code>def getBestScore(target_y, pred_val_y):\n    pred_flattened = pred_val_y.flatten()\n    score_iter_range = np.arange(0.1, 1, 0.01)\n    c = (np.repeat([pred_flattened],len(score_iter_range), axis=0).transpose() &gt; score_iter_range).astype(int).transpose()\n    act_pos = np.bincount(target_y)\n    best_score_value = 0\n    best_score_thresh = 0\n    best_score_precision = 0\n    best_score_recall = 0\n    best_score_pos_acc = 0\n    best_score_neg_acc = 0\n    for x in range(c.shape[0]):\n        p = c[x]\n        tp = target_y == p\n        tp_bins = target_y[tp]\n        equal_counts = np.bincount(tp_bins)\n        if len(equal_counts) &lt; 2:\n            break;\n        tpos = equal_counts[1]\n        tneg = equal_counts[0]\n        if tpos == 0 or tneg == 0:\n            break;\n        c_precision = tpos/ (tpos + act_pos[0] - tneg)\n        c_recall = tpos/ (tpos + act_pos[1] - tpos)\n        score = (c_precision * c_recall * 2)/ (c_precision + c_recall)\n        if score &gt; best_score_value:\n            best_score_value = score\n            best_score_precision = c_precision\n            best_score_recall = c_recall\n            best_score_pos_acc = (equal_counts[1]/act_pos[1])\n            best_score_neg_acc = (equal_counts[0]/act_pos[0])\n            best_score_thresh = score_iter_range[x]\n    return (best_score_value, best_score_precision, best_score_recall, best_score_pos_acc, best_score_neg_acc, np.round(best_score_thresh, 3))\n</code></pre>\n\n<p>P.S: I am not good at python, so any corrections are welcome :)</p>",
  "messages": [
    {
      "id": 464490,
      "postDate": "2019-02-01T01:06:59.707Z",
      "content": "<p>This is a fast implementation to calculate threshold. Instead of calling f1_score manually, this does similar job. </p>\n\n<pre><code>def getBestScore(target_y, pred_val_y):\n    pred_flattened = pred_val_y.flatten()\n    score_iter_range = np.arange(0.1, 1, 0.01)\n    c = (np.repeat([pred_flattened],len(score_iter_range), axis=0).transpose() &gt; score_iter_range).astype(int).transpose()\n    act_pos = np.bincount(target_y)\n    best_score_value = 0\n    best_score_thresh = 0\n    best_score_precision = 0\n    best_score_recall = 0\n    best_score_pos_acc = 0\n    best_score_neg_acc = 0\n    for x in range(c.shape[0]):\n        p = c[x]\n        tp = target_y == p\n        tp_bins = target_y[tp]\n        equal_counts = np.bincount(tp_bins)\n        if len(equal_counts) &lt; 2:\n            break;\n        tpos = equal_counts[1]\n        tneg = equal_counts[0]\n        if tpos == 0 or tneg == 0:\n            break;\n        c_precision = tpos/ (tpos + act_pos[0] - tneg)\n        c_recall = tpos/ (tpos + act_pos[1] - tpos)\n        score = (c_precision * c_recall * 2)/ (c_precision + c_recall)\n        if score &gt; best_score_value:\n            best_score_value = score\n            best_score_precision = c_precision\n            best_score_recall = c_recall\n            best_score_pos_acc = (equal_counts[1]/act_pos[1])\n            best_score_neg_acc = (equal_counts[0]/act_pos[0])\n            best_score_thresh = score_iter_range[x]\n    return (best_score_value, best_score_precision, best_score_recall, best_score_pos_acc, best_score_neg_acc, np.round(best_score_thresh, 3))\n</code></pre>\n\n<p>P.S: I am not good at python, so any corrections are welcome :)</p>",
      "rawMarkdown": "This is a fast implementation to calculate threshold. Instead of calling f1_score manually, this does similar job. \n\n    def getBestScore(target_y, pred_val_y):\n        pred_flattened = pred_val_y.flatten()\n        score_iter_range = np.arange(0.1, 1, 0.01)\n        c = (np.repeat([pred_flattened],len(score_iter_range), axis=0).transpose() &gt; score_iter_range).astype(int).transpose()\n        act_pos = np.bincount(target_y)\n        best_score_value = 0\n        best_score_thresh = 0\n        best_score_precision = 0\n        best_score_recall = 0\n        best_score_pos_acc = 0\n        best_score_neg_acc = 0\n        for x in range(c.shape[0]):\n            p = c[x]\n            tp = target_y == p\n            tp_bins = target_y[tp]\n            equal_counts = np.bincount(tp_bins)\n            if len(equal_counts) &lt; 2:\n                break;\n            tpos = equal_counts[1]\n            tneg = equal_counts[0]\n            if tpos == 0 or tneg == 0:\n                break;\n            c_precision = tpos/ (tpos + act_pos[0] - tneg)\n            c_recall = tpos/ (tpos + act_pos[1] - tpos)\n            score = (c_precision * c_recall * 2)/ (c_precision + c_recall)\n            if score &gt; best_score_value:\n                best_score_value = score\n                best_score_precision = c_precision\n                best_score_recall = c_recall\n                best_score_pos_acc = (equal_counts[1]/act_pos[1])\n                best_score_neg_acc = (equal_counts[0]/act_pos[0])\n                best_score_thresh = score_iter_range[x]\n        return (best_score_value, best_score_precision, best_score_recall, best_score_pos_acc, best_score_neg_acc, np.round(best_score_thresh, 3))\n\n\nP.S: I am not good at python, so any corrections are welcome :)",
      "votes": -1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "464490": "This is a fast implementation to calculate threshold. Instead of calling f1_score manually, this does similar job. \n\n    def getBestScore(target_y, pred_val_y):\n        pred_flattened = pred_val_y.flatten()\n        score_iter_range = np.arange(0.1, 1, 0.01)\n        c = (np.repeat([pred_flattened],len(score_iter_range), axis=0).transpose() &gt; score_iter_range).astype(int).transpose()\n        act_pos = np.bincount(target_y)\n        best_score_value = 0\n        best_score_thresh = 0\n        best_score_precision = 0\n        best_score_recall = 0\n        best_score_pos_acc = 0\n        best_score_neg_acc = 0\n        for x in range(c.shape[0]):\n            p = c[x]\n            tp = target_y == p\n            tp_bins = target_y[tp]\n            equal_counts = np.bincount(tp_bins)\n            if len(equal_counts) &lt; 2:\n                break;\n            tpos = equal_counts[1]\n            tneg = equal_counts[0]\n            if tpos == 0 or tneg == 0:\n                break;\n            c_precision = tpos/ (tpos + act_pos[0] - tneg)\n            c_recall = tpos/ (tpos + act_pos[1] - tpos)\n            score = (c_precision * c_recall * 2)/ (c_precision + c_recall)\n            if score &gt; best_score_value:\n                best_score_value = score\n                best_score_precision = c_precision\n                best_score_recall = c_recall\n                best_score_pos_acc = (equal_counts[1]/act_pos[1])\n                best_score_neg_acc = (equal_counts[0]/act_pos[0])\n                best_score_thresh = score_iter_range[x]\n        return (best_score_value, best_score_precision, best_score_recall, best_score_pos_acc, best_score_neg_acc, np.round(best_score_thresh, 3))\n\n\nP.S: I am not good at python, so any corrections are welcome :)"
  }
}