{
  "id": 208031,
  "title": "15X Faster Precise AUC Calculation",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208031",
  "author_name": "",
  "post_date": "2021-01-01T13:27:49.164745800Z",
  "votes": 31,
  "comment_count": 9,
  "views": 0,
  "content": "<p>A faster way to calculate AUC is at least essential in the following two cases:</p>\n<ul>\n<li>Training the model for many epochs and need to calculate the AUC on both training and test set for many times.</li>\n<li>Use hyper-parameters tuning to get the best weights for ensemble models (my need to loop thousands of times to get optimal weights)</li>\n</ul>\n<p>So, inspired by this <a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013\" target=\"_blank\">discussion</a>, I implemented a faster AUC calculation function with <code>jit</code> support, which is 15X times faster and get the same result as <code>sklearn.metrics.roc_auc_score</code>:</p>\n<h2>Code</h2>\n<pre><code>from numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef auc(actual, predicted):\n    pred_ranks = rankdata(predicted)\n    return _auc(actual, pred_ranks)\n</code></pre>\n<h2>Tests</h2>\n<pre><code>from sklearn.metrics import roc_auc_score\n\ny_true = np.array([1,1,0,0,1,1,0])  \ny_scores = np.array([0.8,0.7,0.5,0.5,0.5,0.5,0.3])\n\nauc(y_true, y_scores)\n#Out: 0.8333333333333334\n\nroc_auc_score(y_true, y_scores)\n#Out: 0.8333333333333334\n\n%%timeit\nauc(y_true, y_scores)\n# Out: 45.8 µs ± 949 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n\n%%timeit\nroc_auc_score(y_true, y_scores)\n683 µs ± 6.28 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)\n</code></pre>",
  "messages": [
    {
      "id": "1134666",
      "postDate": "01/01/2021 13:27:49",
      "content": "<p>A faster way to calculate AUC is at least essential in the following two cases:</p>\n<ul>\n<li>Training the model for many epochs and need to calculate the AUC on both training and test set for many times.</li>\n<li>Use hyper-parameters tuning to get the best weights for ensemble models (my need to loop thousands of times to get optimal weights)</li>\n</ul>\n<p>So, inspired by this <a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013\" target=\"_blank\">discussion</a>, I implemented a faster AUC calculation function with <code>jit</code> support, which is 15X times faster and get the same result as <code>sklearn.metrics.roc_auc_score</code>:</p>\n<h2>Code</h2>\n<pre><code>from numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef auc(actual, predicted):\n    pred_ranks = rankdata(predicted)\n    return _auc(actual, pred_ranks)\n</code></pre>\n<h2>Tests</h2>\n<pre><code>from sklearn.metrics import roc_auc_score\n\ny_true = np.array([1,1,0,0,1,1,0])  \ny_scores = np.array([0.8,0.7,0.5,0.5,0.5,0.5,0.3])\n\nauc(y_true, y_scores)\n#Out: 0.8333333333333334\n\nroc_auc_score(y_true, y_scores)\n#Out: 0.8333333333333334\n\n%%timeit\nauc(y_true, y_scores)\n# Out: 45.8 µs ± 949 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n\n%%timeit\nroc_auc_score(y_true, y_scores)\n683 µs ± 6.28 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)\n</code></pre>",
      "rawMarkdown": "A faster way to calculate AUC is at least essential in the following two cases:\n* Training the model for many epochs and need to calculate the AUC on both training and test set for many times.\n* Use hyper-parameters tuning to get the best weights for ensemble models (my need to loop thousands of times to get optimal weights)\n\nSo, inspired by this [discussion](https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013), I implemented a faster AUC calculation function with `jit` support, which is 15X times faster and get the same result as `sklearn.metrics.roc_auc_score`:\n\n## Code\n\n```Python\nfrom numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef auc(actual, predicted):\n    pred_ranks = rankdata(predicted)\n    return _auc(actual, pred_ranks)\n```\n\n## Tests\n```Python\nfrom sklearn.metrics import roc_auc_score\n\ny_true = np.array([1,1,0,0,1,1,0])  \ny_scores = np.array([0.8,0.7,0.5,0.5,0.5,0.5,0.3])\n\nauc(y_true, y_scores)\n#Out: 0.8333333333333334\n\nroc_auc_score(y_true, y_scores)\n#Out: 0.8333333333333334\n\n%%timeit\nauc(y_true, y_scores)\n# Out: 45.8 µs ± 949 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n\n%%timeit\nroc_auc_score(y_true, y_scores)\n683 µs ± 6.28 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)\n```",
      "votes": null
    },
    {
      "id": "1134774",
      "postDate": "01/01/2021 15:13:26",
      "content": "<p>Money in the bank.</p>",
      "rawMarkdown": "Money in the bank.",
      "votes": null
    },
    {
      "id": "1134797",
      "postDate": "01/01/2021 15:30:07",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>, just curious, what does \"Money in the bank\" mean?</p>",
      "rawMarkdown": "Hi @authman, just curious, what does \"Money in the bank\" mean?",
      "votes": null
    },
    {
      "id": "1134817",
      "postDate": "01/01/2021 15:48:29",
      "content": "<p>Ah, sorry it's an <a href=\"https://www.englishforums.com/English/MoneyInTheBank/jzclc/post.htm\" target=\"_blank\">catch-phrase</a> that just means like \"easy win\" or \"for sure\" or \"instant profit/benefit\". Thank you for sharing the code.</p>",
      "rawMarkdown": "Ah, sorry it's an [catch-phrase](https://www.englishforums.com/English/MoneyInTheBank/jzclc/post.htm) that just means like \"easy win\" or \"for sure\" or \"instant profit/benefit\". Thank you for sharing the code.",
      "votes": null
    },
    {
      "id": "1135361",
      "postDate": "01/02/2021 07:27:54",
      "content": "<p>Thanks for sharing, I  made a little modification to your code to make it run 19x faster than the original sklearn implementation.</p>\n<pre><code>from numba import njit,jit\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = 0\n    for k in actual:\n        n_pos += k\n    n_neg = len(actual) - n_pos\n    nominator = 0\n\n    for idx,val in enumerate(actual):\n        if val==1:\n            nominator += pred_ranks[idx]\n\n    return (nominator - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\n\ndef rank_data(x):\n    #Inspired from : https://stackoverflow.com/questions/14671013/ranking-of-numpy-array-with-possible-duplicates\n    u, v = np.unique(x, return_inverse=True)\n    up = (np.cumsum(np.bincount(v)) - 1)[v] +1\n    low = (np.cumsum(np.concatenate(([0], np.bincount(v)))))[v] +1\n    return (up+low)/2\n\ndef auc(actual, predicted):\n    pred_ranks = rank_data(predicted)\n    return _auc(actual, pred_ranks)\n\n%%timeit\nauc(y_true,y_scores)`\n35.3 µs ± 1.01 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n</code></pre>",
      "rawMarkdown": "Thanks for sharing, I  made a little modification to your code to make it run 19x faster than the original sklearn implementation.\n\n\n```\nfrom numba import njit,jit\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = 0\n    for k in actual:\n        n_pos += k\n    n_neg = len(actual) - n_pos\n    nominator = 0\n    \n    for idx,val in enumerate(actual):\n        if val==1:\n            nominator += pred_ranks[idx]\n            \n    return (nominator - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\n\ndef rank_data(x):\n    #Inspired from : https://stackoverflow.com/questions/14671013/ranking-of-numpy-array-with-possible-duplicates\n    u, v = np.unique(x, return_inverse=True)\n    up = (np.cumsum(np.bincount(v)) - 1)[v] +1\n    low = (np.cumsum(np.concatenate(([0], np.bincount(v)))))[v] +1\n    return (up+low)/2\n    \ndef auc(actual, predicted):\n    pred_ranks = rank_data(predicted)\n    return _auc(actual, pred_ranks)\n\n%%timeit\nauc(y_true,y_scores)`\n35.3 µs ± 1.01 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n```",
      "votes": null
    },
    {
      "id": "1135363",
      "postDate": "01/02/2021 07:33:33",
      "content": "<p>But after further investigation, our method doesn't scale too well on a large dataset (array with length 1 million)<br>\nIn large array, sklearn took 394ms and  your method(and mine) took around 220ms 😅</p>",
      "rawMarkdown": "But after further investigation, our method doesn't scale too well on a large dataset (array with length 1 million)\nIn large array, sklearn took 394ms and  your method(and mine) took around 220ms 😅",
      "votes": null
    },
    {
      "id": "1135395",
      "postDate": "01/02/2021 08:06:57",
      "content": "<p>Thank you for sharing this. My question is:  have you tried to use this code in lgbm training process, is it faster than using the default auc metric?<br>\nI tried the below code, but it seems longer than just using default auc metric</p>\n<pre><code>from numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef fast_auc( predicted,actual):\n    pred_ranks = rankdata(predicted)\n    return \"f_auc\", _auc(actual.get_label(), pred_ranks), True\n\nmodel = lgb.train(\n                    {'objective': 'binary', \n                    'seed':2020,\"learning_rate\":0.05}, \n                    lgb_train,\n                    valid_sets=[lgb_valid],\n                    evals_result=evals_result,\n                    feval=fast_auc,\n                    verbose_eval=100,\n                    num_boost_round=10000,\n                    early_stopping_rounds=100\n                )\n</code></pre>",
      "rawMarkdown": "Thank you for sharing this. My question is:  have you tried to use this code in lgbm training process, is it faster than using the default auc metric?\nI tried the below code, but it seems longer than just using default auc metric\n```\nfrom numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef fast_auc( predicted,actual):\n    pred_ranks = rankdata(predicted)\n    return \"f_auc\", _auc(actual.get_label(), pred_ranks), True\n\nmodel = lgb.train(\n                    {'objective': 'binary', \n                    'seed':2020,\"learning_rate\":0.05}, \n                    lgb_train,\n                    valid_sets=[lgb_valid],\n                    evals_result=evals_result,\n                    feval=fast_auc,\n                    verbose_eval=100,\n                    num_boost_round=10000,\n                    early_stopping_rounds=100\n                )\n```",
      "votes": null
    },
    {
      "id": "1136346",
      "postDate": "01/03/2021 01:40:56",
      "content": "<p>Thanks for sharing! Although useful, the bottleneck for me usually is the model making its predictions on the evaluation set and not on the <em>calculation</em> of the evaluation metric. Numba is just amazing, being able to so effortlessly speed up codes!</p>",
      "rawMarkdown": "Thanks for sharing! Although useful, the bottleneck for me usually is the model making its predictions on the evaluation set and not on the *calculation* of the evaluation metric. Numba is just amazing, being able to so effortlessly speed up codes!",
      "votes": null
    },
    {
      "id": "1136368",
      "postDate": "01/03/2021 02:22:28",
      "content": "<p>Take a look at <a href=\"https://github.com/microsoft/LightGBM/blob/3c0e12dc5cf1bdb941e9845d08ba7b891a935745/src/metric/binary_metric.hpp\" target=\"_blank\">lgbm implementation of auc</a><br>\nThey use native C++ , no wonder their implementation is faster 😅</p>",
      "rawMarkdown": "Take a look at [lgbm implementation of auc](https://github.com/microsoft/LightGBM/blob/3c0e12dc5cf1bdb941e9845d08ba7b891a935745/src/metric/binary_metric.hpp)\nThey use native C++ , no wonder their implementation is faster 😅",
      "votes": null
    },
    {
      "id": "1146601",
      "postDate": "01/09/2021 21:53:57",
      "content": "<p>I think a better way is to request the <code>LGB</code> community to provide Python APIs for the calculations of metrics/losses implemented in <code>LGB</code>. I created an <a href=\"https://github.com/microsoft/LightGBM/issues/3744\" target=\"_blank\">issue</a> for this purpose.</p>",
      "rawMarkdown": "I think a better way is to request the `LGB` community to provide Python APIs for the calculations of metrics/losses implemented in `LGB`. I created an [issue](https://github.com/microsoft/LightGBM/issues/3744) for this purpose.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1134774,
      "author_name": "authman",
      "author_url": "",
      "post_date": "01/01/2021 15:13:26",
      "content": "<p>Money in the bank.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1134797,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/01/2021 15:30:07",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>, just curious, what does \"Money in the bank\" mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134817,
          "author_name": "authman",
          "author_url": "",
          "post_date": "01/01/2021 15:48:29",
          "content": "<p>Ah, sorry it's an <a href=\"https://www.englishforums.com/English/MoneyInTheBank/jzclc/post.htm\" target=\"_blank\">catch-phrase</a> that just means like \"easy win\" or \"for sure\" or \"instant profit/benefit\". Thank you for sharing the code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1135361,
      "author_name": "vinson2233",
      "author_url": "",
      "post_date": "01/02/2021 07:27:54",
      "content": "<p>Thanks for sharing, I  made a little modification to your code to make it run 19x faster than the original sklearn implementation.</p>\n<pre><code>from numba import njit,jit\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = 0\n    for k in actual:\n        n_pos += k\n    n_neg = len(actual) - n_pos\n    nominator = 0\n\n    for idx,val in enumerate(actual):\n        if val==1:\n            nominator += pred_ranks[idx]\n\n    return (nominator - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\n\ndef rank_data(x):\n    #Inspired from : https://stackoverflow.com/questions/14671013/ranking-of-numpy-array-with-possible-duplicates\n    u, v = np.unique(x, return_inverse=True)\n    up = (np.cumsum(np.bincount(v)) - 1)[v] +1\n    low = (np.cumsum(np.concatenate(([0], np.bincount(v)))))[v] +1\n    return (up+low)/2\n\ndef auc(actual, predicted):\n    pred_ranks = rank_data(predicted)\n    return _auc(actual, pred_ranks)\n\n%%timeit\nauc(y_true,y_scores)`\n35.3 µs ± 1.01 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1135363,
          "author_name": "vinson2233",
          "author_url": "",
          "post_date": "01/02/2021 07:33:33",
          "content": "<p>But after further investigation, our method doesn't scale too well on a large dataset (array with length 1 million)<br>\nIn large array, sklearn took 394ms and  your method(and mine) took around 220ms 😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1135395,
      "author_name": "kaibochen",
      "author_url": "",
      "post_date": "01/02/2021 08:06:57",
      "content": "<p>Thank you for sharing this. My question is:  have you tried to use this code in lgbm training process, is it faster than using the default auc metric?<br>\nI tried the below code, but it seems longer than just using default auc metric</p>\n<pre><code>from numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef fast_auc( predicted,actual):\n    pred_ranks = rankdata(predicted)\n    return \"f_auc\", _auc(actual.get_label(), pred_ranks), True\n\nmodel = lgb.train(\n                    {'objective': 'binary', \n                    'seed':2020,\"learning_rate\":0.05}, \n                    lgb_train,\n                    valid_sets=[lgb_valid],\n                    evals_result=evals_result,\n                    feval=fast_auc,\n                    verbose_eval=100,\n                    num_boost_round=10000,\n                    early_stopping_rounds=100\n                )\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1136368,
          "author_name": "vinson2233",
          "author_url": "",
          "post_date": "01/03/2021 02:22:28",
          "content": "<p>Take a look at <a href=\"https://github.com/microsoft/LightGBM/blob/3c0e12dc5cf1bdb941e9845d08ba7b891a935745/src/metric/binary_metric.hpp\" target=\"_blank\">lgbm implementation of auc</a><br>\nThey use native C++ , no wonder their implementation is faster 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146601,
          "author_name": "wuwenmin",
          "author_url": "",
          "post_date": "01/09/2021 21:53:57",
          "content": "<p>I think a better way is to request the <code>LGB</code> community to provide Python APIs for the calculations of metrics/losses implemented in <code>LGB</code>. I created an <a href=\"https://github.com/microsoft/LightGBM/issues/3744\" target=\"_blank\">issue</a> for this purpose.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136346,
      "author_name": "doctorkael",
      "author_url": "",
      "post_date": "01/03/2021 01:40:56",
      "content": "<p>Thanks for sharing! Although useful, the bottleneck for me usually is the model making its predictions on the evaluation set and not on the <em>calculation</em> of the evaluation metric. Numba is just amazing, being able to so effortlessly speed up codes!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1134666": "A faster way to calculate AUC is at least essential in the following two cases:\n* Training the model for many epochs and need to calculate the AUC on both training and test set for many times.\n* Use hyper-parameters tuning to get the best weights for ensemble models (my need to loop thousands of times to get optimal weights)\n\nSo, inspired by this [discussion](https://www.kaggle.com/c/microsoft-malware-prediction/discussion/76013), I implemented a faster AUC calculation function with `jit` support, which is 15X times faster and get the same result as `sklearn.metrics.roc_auc_score`:\n\n## Code\n\n```Python\nfrom numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef auc(actual, predicted):\n    pred_ranks = rankdata(predicted)\n    return _auc(actual, pred_ranks)\n```\n\n## Tests\n```Python\nfrom sklearn.metrics import roc_auc_score\n\ny_true = np.array([1,1,0,0,1,1,0])  \ny_scores = np.array([0.8,0.7,0.5,0.5,0.5,0.5,0.3])\n\nauc(y_true, y_scores)\n#Out: 0.8333333333333334\n\nroc_auc_score(y_true, y_scores)\n#Out: 0.8333333333333334\n\n%%timeit\nauc(y_true, y_scores)\n# Out: 45.8 µs ± 949 ns per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n\n%%timeit\nroc_auc_score(y_true, y_scores)\n683 µs ± 6.28 µs per loop (mean ± std. dev. of 7 runs, 1000 loops each)\n```",
    "1134774": "Money in the bank.",
    "1134797": "Hi @authman, just curious, what does \"Money in the bank\" mean?",
    "1134817": "Ah, sorry it's an [catch-phrase](https://www.englishforums.com/English/MoneyInTheBank/jzclc/post.htm) that just means like \"easy win\" or \"for sure\" or \"instant profit/benefit\". Thank you for sharing the code.",
    "1135361": "Thanks for sharing, I  made a little modification to your code to make it run 19x faster than the original sklearn implementation.\n\n\n```\nfrom numba import njit,jit\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = 0\n    for k in actual:\n        n_pos += k\n    n_neg = len(actual) - n_pos\n    nominator = 0\n    \n    for idx,val in enumerate(actual):\n        if val==1:\n            nominator += pred_ranks[idx]\n            \n    return (nominator - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\n\ndef rank_data(x):\n    #Inspired from : https://stackoverflow.com/questions/14671013/ranking-of-numpy-array-with-possible-duplicates\n    u, v = np.unique(x, return_inverse=True)\n    up = (np.cumsum(np.bincount(v)) - 1)[v] +1\n    low = (np.cumsum(np.concatenate(([0], np.bincount(v)))))[v] +1\n    return (up+low)/2\n    \ndef auc(actual, predicted):\n    pred_ranks = rank_data(predicted)\n    return _auc(actual, pred_ranks)\n\n%%timeit\nauc(y_true,y_scores)`\n35.3 µs ± 1.01 µs per loop (mean ± std. dev. of 7 runs, 10000 loops each)\n```",
    "1135363": "But after further investigation, our method doesn't scale too well on a large dataset (array with length 1 million)\nIn large array, sklearn took 394ms and  your method(and mine) took around 220ms 😅",
    "1135395": "Thank you for sharing this. My question is:  have you tried to use this code in lgbm training process, is it faster than using the default auc metric?\nI tried the below code, but it seems longer than just using default auc metric\n```\nfrom numba import njit\nfrom scipy.stats import rankdata\n\n@njit\ndef _auc(actual, pred_ranks):\n    actual = np.asarray(actual)\n    pred_ranks = np.asarray(pred_ranks)\n    n_pos = np.sum(actual)\n    n_neg = len(actual) - n_pos\n    return (np.sum(pred_ranks[actual==1]) - n_pos*(n_pos+1)/2) / (n_pos*n_neg)\n\ndef fast_auc( predicted,actual):\n    pred_ranks = rankdata(predicted)\n    return \"f_auc\", _auc(actual.get_label(), pred_ranks), True\n\nmodel = lgb.train(\n                    {'objective': 'binary', \n                    'seed':2020,\"learning_rate\":0.05}, \n                    lgb_train,\n                    valid_sets=[lgb_valid],\n                    evals_result=evals_result,\n                    feval=fast_auc,\n                    verbose_eval=100,\n                    num_boost_round=10000,\n                    early_stopping_rounds=100\n                )\n```",
    "1136346": "Thanks for sharing! Although useful, the bottleneck for me usually is the model making its predictions on the evaluation set and not on the *calculation* of the evaluation metric. Numba is just amazing, being able to so effortlessly speed up codes!",
    "1136368": "Take a look at [lgbm implementation of auc](https://github.com/microsoft/LightGBM/blob/3c0e12dc5cf1bdb941e9845d08ba7b891a935745/src/metric/binary_metric.hpp)\nThey use native C++ , no wonder their implementation is faster 😅",
    "1146601": "I think a better way is to request the `LGB` community to provide Python APIs for the calculations of metrics/losses implemented in `LGB`. I created an [issue](https://github.com/microsoft/LightGBM/issues/3744) for this purpose."
  },
  "source": "meta"
}