{
  "id": 75949,
  "title": "yet another rounder for f1",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75949",
  "author_name": "",
  "post_date": "2018-12-27T20:08:38.458410Z",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So, here is yet another rounder for f1 which uses nelder-mead optimization.</p>\n\n<pre><code>from sklearn import metrics\nfrom functools import partial\nimport scipy as sp\n\nclass f1Rounder(object):\n    def __init__(self):\n        self.coef_ = [0]\n\n    def _f1_loss(self, coef, X, y):\n        X_p = np.copy(X)\n        for i, pred in enumerate(X_p):\n            if pred &lt; coef[0]:\n                X_p[i] = 0\n            else:\n                X_p[i] = 1\n\n        ll = metrics.f1_score(y, X_p)\n        return -ll\n\n    def fit(self, X, y):\n        loss_partial = partial(self._f1_loss, X=X, y=y)\n        initial_coef = [0.5]\n        self.coef_ = sp.optimize.minimize(loss_partial, initial_coef, method='nelder-mead')\n\n    def predict(self, X):\n        X_p = np.copy(X)\n        print(self.coef_['x'])\n        coef = self.coef_['x']\n        for i, pred in enumerate(X_p):\n            if pred &lt; coef[0]:\n                X_p[i] = 0\n            else:\n                X_p[i] = 1\n        return X_p\n</code></pre>\n\n<p>Hope some of you find it useful ;)</p>",
  "messages": [
    {
      "id": "446289",
      "postDate": "12/27/2018 20:08:38",
      "content": "<p>So, here is yet another rounder for f1 which uses nelder-mead optimization.</p>\n\n<pre><code>from sklearn import metrics\nfrom functools import partial\nimport scipy as sp\n\nclass f1Rounder(object):\n    def __init__(self):\n        self.coef_ = [0]\n\n    def _f1_loss(self, coef, X, y):\n        X_p = np.copy(X)\n        for i, pred in enumerate(X_p):\n            if pred &lt; coef[0]:\n                X_p[i] = 0\n            else:\n                X_p[i] = 1\n\n        ll = metrics.f1_score(y, X_p)\n        return -ll\n\n    def fit(self, X, y):\n        loss_partial = partial(self._f1_loss, X=X, y=y)\n        initial_coef = [0.5]\n        self.coef_ = sp.optimize.minimize(loss_partial, initial_coef, method='nelder-mead')\n\n    def predict(self, X):\n        X_p = np.copy(X)\n        print(self.coef_['x'])\n        coef = self.coef_['x']\n        for i, pred in enumerate(X_p):\n            if pred &lt; coef[0]:\n                X_p[i] = 0\n            else:\n                X_p[i] = 1\n        return X_p\n</code></pre>\n\n<p>Hope some of you find it useful ;)</p>",
      "rawMarkdown": "So, here is yet another rounder for f1 which uses nelder-mead optimization.\n\n\n\n    from sklearn import metrics\n    from functools import partial\n    import scipy as sp\n    \n    class f1Rounder(object):\n        def __init__(self):\n            self.coef_ = [0]\n    \n        def _f1_loss(self, coef, X, y):\n            X_p = np.copy(X)\n            for i, pred in enumerate(X_p):\n                if pred &lt; coef[0]:\n                    X_p[i] = 0\n                else:\n                    X_p[i] = 1\n    \n            ll = metrics.f1_score(y, X_p)\n            return -ll\n    \n        def fit(self, X, y):\n            loss_partial = partial(self._f1_loss, X=X, y=y)\n            initial_coef = [0.5]\n            self.coef_ = sp.optimize.minimize(loss_partial, initial_coef, method='nelder-mead')\n    \n        def predict(self, X):\n            X_p = np.copy(X)\n            print(self.coef_['x'])\n            coef = self.coef_['x']\n            for i, pred in enumerate(X_p):\n                if pred &lt; coef[0]:\n                    X_p[i] = 0\n                else:\n                    X_p[i] = 1\n            return X_p\n\nHope some of you find it useful ;)",
      "votes": null
    },
    {
      "id": "446292",
      "postDate": "12/27/2018 20:11:32",
      "content": "<p>inspired/taken from old kaggle competitions. :)</p>",
      "rawMarkdown": "inspired/taken from old kaggle competitions. :)",
      "votes": null
    },
    {
      "id": "446497",
      "postDate": "12/28/2018 07:31:23",
      "content": "<p>I don't fully understand, how would you use this? as a metric for sklearn classifiers?</p>",
      "rawMarkdown": "I don't fully understand, how would you use this? as a metric for sklearn classifiers?",
      "votes": null
    },
    {
      "id": "446629",
      "postDate": "12/28/2018 12:05:08",
      "content": "<p>once you have kfold predictions, you can do:</p>\n\n<p>f1r = f1Rounder()</p>\n\n<p>f1r.fit(train_kfold, y)</p>\n\n<p>preds = f1r.predict(test_predictions)</p>",
      "rawMarkdown": "once you have kfold predictions, you can do:\n\n\nf1r = f1Rounder()\n\nf1r.fit(train_kfold, y)\n\npreds = f1r.predict(test_predictions)",
      "votes": null
    },
    {
      "id": "446633",
      "postDate": "12/28/2018 12:12:22",
      "content": "<p>So I guess the aim of this is to find the best threshold without the need to search among all possible thresholds?</p>",
      "rawMarkdown": "So I guess the aim of this is to find the best threshold without the need to search among all possible thresholds?",
      "votes": null
    },
    {
      "id": "446861",
      "postDate": "12/28/2018 19:50:28",
      "content": "<p>yep :)</p>",
      "rawMarkdown": "yep :)",
      "votes": null
    },
    {
      "id": "447905",
      "postDate": "12/30/2018 19:58:08",
      "content": "<p>I used the method and got same score with the following code.</p>\n\n<pre><code>def f1_smart(y_true, y_pred):\n\n    args = np.argsort(y_pred)\n\n    tp = y_true.sum()\n\n    fs = (tp - np.cumsum(y_true[args[:-1]])) / np.arange(y_true.shape[0] + tp - 1, tp, -1)\n\n    res_idx = np.argmax(fs)\n\n    return 2 * fs[res_idx], (y_pred[args[res_idx]] + y_pred[args[res_idx + 1]]) / 2\n</code></pre>\n\n<p>And, <code>f1Rounder</code> is slower.</p>",
      "rawMarkdown": "I used the method and got same score with the following code.\n\n    def f1_smart(y_true, y_pred):\n\n        args = np.argsort(y_pred)\n\n        tp = y_true.sum()\n\n        fs = (tp - np.cumsum(y_true[args[:-1]])) / np.arange(y_true.shape[0] + tp - 1, tp, -1)\n\n        res_idx = np.argmax(fs)\n\n        return 2 * fs[res_idx], (y_pred[args[res_idx]] + y_pred[args[res_idx + 1]]) / 2\n\nAnd, `f1Rounder` is slower.",
      "votes": null
    },
    {
      "id": "465937",
      "postDate": "02/04/2019 11:07:41",
      "content": "<p>can you explain how this works ? </p>",
      "rawMarkdown": "can you explain how this works ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 446292,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "12/27/2018 20:11:32",
      "content": "<p>inspired/taken from old kaggle competitions. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 446497,
      "author_name": "david26694",
      "author_url": "",
      "post_date": "12/28/2018 07:31:23",
      "content": "<p>I don't fully understand, how would you use this? as a metric for sklearn classifiers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 446629,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "12/28/2018 12:05:08",
          "content": "<p>once you have kfold predictions, you can do:</p>\n\n<p>f1r = f1Rounder()</p>\n\n<p>f1r.fit(train_kfold, y)</p>\n\n<p>preds = f1r.predict(test_predictions)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446633,
          "author_name": "david26694",
          "author_url": "",
          "post_date": "12/28/2018 12:12:22",
          "content": "<p>So I guess the aim of this is to find the best threshold without the need to search among all possible thresholds?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446861,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "12/28/2018 19:50:28",
          "content": "<p>yep :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447905,
      "author_name": "syhens",
      "author_url": "",
      "post_date": "12/30/2018 19:58:08",
      "content": "<p>I used the method and got same score with the following code.</p>\n\n<pre><code>def f1_smart(y_true, y_pred):\n\n    args = np.argsort(y_pred)\n\n    tp = y_true.sum()\n\n    fs = (tp - np.cumsum(y_true[args[:-1]])) / np.arange(y_true.shape[0] + tp - 1, tp, -1)\n\n    res_idx = np.argmax(fs)\n\n    return 2 * fs[res_idx], (y_pred[args[res_idx]] + y_pred[args[res_idx + 1]]) / 2\n</code></pre>\n\n<p>And, <code>f1Rounder</code> is slower.</p>",
      "votes": null,
      "replies": [
        {
          "id": 465937,
          "author_name": "vonneumann",
          "author_url": "",
          "post_date": "02/04/2019 11:07:41",
          "content": "<p>can you explain how this works ? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "446289": "So, here is yet another rounder for f1 which uses nelder-mead optimization.\n\n\n\n    from sklearn import metrics\n    from functools import partial\n    import scipy as sp\n    \n    class f1Rounder(object):\n        def __init__(self):\n            self.coef_ = [0]\n    \n        def _f1_loss(self, coef, X, y):\n            X_p = np.copy(X)\n            for i, pred in enumerate(X_p):\n                if pred &lt; coef[0]:\n                    X_p[i] = 0\n                else:\n                    X_p[i] = 1\n    \n            ll = metrics.f1_score(y, X_p)\n            return -ll\n    \n        def fit(self, X, y):\n            loss_partial = partial(self._f1_loss, X=X, y=y)\n            initial_coef = [0.5]\n            self.coef_ = sp.optimize.minimize(loss_partial, initial_coef, method='nelder-mead')\n    \n        def predict(self, X):\n            X_p = np.copy(X)\n            print(self.coef_['x'])\n            coef = self.coef_['x']\n            for i, pred in enumerate(X_p):\n                if pred &lt; coef[0]:\n                    X_p[i] = 0\n                else:\n                    X_p[i] = 1\n            return X_p\n\nHope some of you find it useful ;)",
    "446292": "inspired/taken from old kaggle competitions. :)",
    "446497": "I don't fully understand, how would you use this? as a metric for sklearn classifiers?",
    "446629": "once you have kfold predictions, you can do:\n\n\nf1r = f1Rounder()\n\nf1r.fit(train_kfold, y)\n\npreds = f1r.predict(test_predictions)",
    "446633": "So I guess the aim of this is to find the best threshold without the need to search among all possible thresholds?",
    "446861": "yep :)",
    "447905": "I used the method and got same score with the following code.\n\n    def f1_smart(y_true, y_pred):\n\n        args = np.argsort(y_pred)\n\n        tp = y_true.sum()\n\n        fs = (tp - np.cumsum(y_true[args[:-1]])) / np.arange(y_true.shape[0] + tp - 1, tp, -1)\n\n        res_idx = np.argmax(fs)\n\n        return 2 * fs[res_idx], (y_pred[args[res_idx]] + y_pred[args[res_idx + 1]]) / 2\n\nAnd, `f1Rounder` is slower.",
    "465937": "can you explain how this works ?"
  },
  "source": "meta"
}