{
  "id": 244387,
  "title": "Fast metric code",
  "url": "/competitions/seti-breakthrough-listen/discussion/244387",
  "author_name": "CPMP",
  "post_date": "2021-06-06T14:43:22.165000",
  "votes": 52,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Here is a code that computes roc-auc.  It is way faster than scikit-learn implementation.  However, if you have ties in your predictions then the value might differ from that of scikit-learn.</p>\n<p>Make sure you respect the order of parameters: the first one must be the target, the second one the predictions.  You can use logits, there is no need to input probability as predictions.  Indeed, what matters is the ordering of your predictions, not their absolute values.</p>\n<pre><code>import numpy as np \nfrom numba import jit\n\n@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n</code></pre>\n<p>Edt.  <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> improved over it using numpy cumsum, see comments.  I include his code here for convenience.</p>\n<pre><code>def fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>\n<p>Edit2.  This code is even faster:</p>\n<pre><code>@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n</code></pre>\n<p>Edit 3: <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> has shared that using GPU is even faster.  Here is his code:|</p>\n<pre><code>def fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>",
  "messages": [
    {
      "id": 1338552,
      "postDate": "2021-06-06T14:43:22.167Z",
      "content": "<p>Here is a code that computes roc-auc.  It is way faster than scikit-learn implementation.  However, if you have ties in your predictions then the value might differ from that of scikit-learn.</p>\n<p>Make sure you respect the order of parameters: the first one must be the target, the second one the predictions.  You can use logits, there is no need to input probability as predictions.  Indeed, what matters is the ordering of your predictions, not their absolute values.</p>\n<pre><code>import numpy as np \nfrom numba import jit\n\n@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n</code></pre>\n<p>Edt.  <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> improved over it using numpy cumsum, see comments.  I include his code here for convenience.</p>\n<pre><code>def fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>\n<p>Edit2.  This code is even faster:</p>\n<pre><code>@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n</code></pre>\n<p>Edit 3: <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> has shared that using GPU is even faster.  Here is his code:|</p>\n<pre><code>def fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>",
      "rawMarkdown": "Here is a code that computes roc-auc.  It is way faster than scikit-learn implementation.  However, if you have ties in your predictions then the value might differ from that of scikit-learn.\n\nMake sure you respect the order of parameters: the first one must be the target, the second one the predictions.  You can use logits, there is no need to input probability as predictions.  Indeed, what matters is the ordering of your predictions, not their absolute values.\n\n```\nimport numpy as np \nfrom numba import jit\n\n@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n```\n\nEdt.  @nofreewill improved over it using numpy cumsum, see comments.  I include his code here for convenience.\n\n```\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n\n```\n\nEdit2.  This code is even faster:\n\n```\n@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n```\n\nEdit 3: @nofreewill has shared that using GPU is even faster.  Here is his code:|\n\n```\n\ndef fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n\n```\n",
      "votes": 52
    },
    {
      "id": 1338610,
      "postDate": "2021-06-06T15:28:19.087Z",
      "content": "<p>Speeking of speed, there is also np.cumsum:</p>\n<pre><code>%%timeit\nnp.cumsum(np.arange(1000))\n</code></pre>\n<p>4.01 µs ± 14.4 ns</p>\n<pre><code>%%timeit\ns = 0\nfor i in range(1000):\n    s += i\n</code></pre>\n<p>35.4 µs ± 229 ns (EDIT: note that numba's jit decorator speeds up the original loop quite a bit!)</p>\n<p>Code utilizing it:</p>\n<pre><code>def fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>",
      "rawMarkdown": "Speeking of speed, there is also np.cumsum:\n\n```\n%%timeit\nnp.cumsum(np.arange(1000))\n```\n4.01 µs ± 14.4 ns\n\n```\n%%timeit\ns = 0\nfor i in range(1000):\n    s += i\n```\n35.4 µs ± 229 ns (EDIT: note that numba's jit decorator speeds up the original loop quite a bit!)\n\nCode utilizing it:\n```\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n```",
      "votes": 9,
      "replies": [
        {
          "id": 1338877,
          "postDate": "2021-06-06T20:05:39.077Z",
          "content": "<p>Your fast auc is 16% faster indeed.  Thanks for sharing.</p>",
          "rawMarkdown": "Your fast auc is 16% faster indeed.  Thanks for sharing.",
          "votes": 2
        },
        {
          "id": 1338886,
          "postDate": "2021-06-06T20:20:13.153Z",
          "content": "<p>On an array of 10000 elements:<br>\n[cumsum vs for loop]</p>\n<p><strong>Without jit:</strong><br>\n[<strong>575 µs</strong> vs 3.36 ms]</p>\n<p><strong>With jit:</strong><br>\n[727 µs vs 709 µs]</p>\n<p>So the best time is <strong>cumsum without numba's jit</strong> decorator, although not at all that much as those initial timings would suggest.<br>\nI didn't know numba, so thank you for introducing it. It smells like magic :D</p>",
          "rawMarkdown": "On an array of 10000 elements:\n[cumsum vs for loop]\n\n**Without jit:**\n[**575 µs** vs 3.36 ms]\n\n**With jit:**\n[727 µs vs 709 µs]\n\nSo the best time is **cumsum without numba's jit** decorator, although not at all that much as those initial timings would suggest.\nI didn't know numba, so thank you for introducing it. It smells like magic :D",
          "votes": 1
        },
        {
          "id": 1338892,
          "postDate": "2021-06-06T20:37:42.987Z",
          "content": "<p>On an array of 1000 elements, though,<br>\nyour for loop with jit is a little better than cumsum. Both are around 37 us.</p>",
          "rawMarkdown": "On an array of 1000 elements, though,\nyour for loop with jit is a little better than cumsum. Both are around 37 us.",
          "votes": 1
        },
        {
          "id": 1338936,
          "postDate": "2021-06-06T21:47:29.843Z",
          "content": "<p>yeah, I found the same.  cumsum wins as the array gets larger.</p>",
          "rawMarkdown": "yeah, I found the same.  cumsum wins as the array gets larger.\n",
          "votes": 1
        },
        {
          "id": 1339681,
          "postDate": "2021-06-07T11:49:05.323Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> I made numba code run faster than yours (by a small margin):</p>\n<pre><code>@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n</code></pre>",
          "rawMarkdown": "@nofreewill I made numba code run faster than yours (by a small margin):\n\n```\n@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n```",
          "votes": 1
        },
        {
          "id": 1339682,
          "postDate": "2021-06-07T11:52:31.550Z",
          "content": "<p>Challange accepted! :D</p>",
          "rawMarkdown": "Challange accepted! :D",
          "votes": 1
        },
        {
          "id": 1339780,
          "postDate": "2021-06-07T12:59:34.417Z",
          "content": "<p>I have no gpu to try it right now, but I'm hopeful that it's faster.<br>\nTill then, here is the code.</p>\n<pre><code>def fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>\n<p>One should collect their y_probs and y_trues on the gpu as they predict, right away, to bypass additional cpu -&gt; gpu communications.</p>",
          "rawMarkdown": "I have no gpu to try it right now, but I'm hopeful that it's faster.\nTill then, here is the code.\n\n```\ndef fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n```\n\nOne should collect their y_probs and y_trues on the gpu as they predict, right away, to bypass additional cpu -> gpu communications.",
          "votes": 3
        },
        {
          "id": 1340053,
          "postDate": "2021-06-07T16:19:28.657Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I could make it faster on gpu with torch's cumsum</p>\n<p><strong>data</strong></p>\n<pre><code>y_true = np.random.randint(2,size=10000)\ny_prob = np.random.rand(10000)\n</code></pre>\n<p><strong>your most recent version</strong></p>\n<pre><code>@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n</code></pre>\n<pre><code>%%timeit\nfast_auc(y_true, y_prob)\n</code></pre>\n<p>537 µs</p>\n<p><strong>fast_auc_torch</strong> - sending to gpu first</p>\n<pre><code>%%timeit\nfast_auc_torch(torch.tensor(y_true).to(device), torch.tensor(y_prob).to(device))\n</code></pre>\n<p>340 µs</p>\n<p><strong>fast_auc_torch</strong> - data is already on gpu</p>\n<pre><code>y_true_tensor = torch.tensor(y_true).to(device), \ny_prob_tensor = torch.tensor(y_prob).to(device)\n</code></pre>\n<pre><code>%%timeit\nfast_auc_torch(y_true_tensor, y_prob_tensor)\n</code></pre>\n<p>205 µs</p>\n<p>On a 1080Ti</p>",
          "rawMarkdown": "@cpmpml I could make it faster on gpu with torch's cumsum\n\n**data**\n```\ny_true = np.random.randint(2,size=10000)\ny_prob = np.random.rand(10000)\n```\n\n**your most recent version**\n```\n@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n```\n```\n%%timeit\nfast_auc(y_true, y_prob)\n```\n537 µs\n\n**fast_auc_torch** - sending to gpu first\n```\n%%timeit\nfast_auc_torch(torch.tensor(y_true).to(device), torch.tensor(y_prob).to(device))\n```\n340 µs\n\n**fast_auc_torch** - data is already on gpu\n```\ny_true_tensor = torch.tensor(y_true).to(device), \ny_prob_tensor = torch.tensor(y_prob).to(device)\n```\n```\n%%timeit\nfast_auc_torch(y_true_tensor, y_prob_tensor)\n```\n205 µs\n\nOn a 1080Ti",
          "votes": 1
        },
        {
          "id": 1340058,
          "postDate": "2021-06-07T16:25:06.827Z",
          "content": "<p>With 1M number of elements.<br>\nsklearn: <strong>261 ms</strong><br>\nnumpy cumsum: <strong>100 ms</strong><br>\nfast_auc_forloop with jit: <strong>131 ms</strong><br>\nfast_auc with separate fast_auc_aux: <strong>95.2 ms</strong><br>\nfast_auc_torch sending data to gpu: <strong>10 ms</strong><br>\nfast_auc_torch data is already on gpu: <strong>2.83 ms</strong></p>",
          "rawMarkdown": "With 1M number of elements.\nsklearn: **261 ms**\nnumpy cumsum: **100 ms**\nfast_auc_forloop with jit: **131 ms**\nfast_auc with separate fast_auc_aux: **95.2 ms**\nfast_auc_torch sending data to gpu: **10 ms**\nfast_auc_torch data is already on gpu: **2.83 ms**",
          "votes": 1
        },
        {
          "id": 1340063,
          "postDate": "2021-06-07T16:32:07.820Z",
          "content": "<p>sklearn on 10k takes 2.4 ms, so in this competition with a few 10k samples, we can save a second in every couple of hundred epochs … xd</p>",
          "rawMarkdown": "sklearn on 10k takes 2.4 ms, so in this competition with a few 10k samples, we can save a second in every couple of hundred epochs ... xd",
          "votes": 1
        },
        {
          "id": 1341504,
          "postDate": "2021-06-08T17:23:46.510Z",
          "content": "<blockquote>\n  <p>Challange accepted! :D</p>\n</blockquote>\n<p>Let's show the Aliens how fast we can roc-auc them ;)</p>\n<p>P.S. I suggest to use CPU only and seti train dataset shape for benchmarking our implementations. Are you?</p>",
          "rawMarkdown": "> Challange accepted! :D\n\nLet's show the Aliens how fast we can roc-auc them ;)\n\nP.S. I suggest to use CPU only and seti train dataset shape for benchmarking our implementations. Are you?\n",
          "votes": 1
        },
        {
          "id": 1342947,
          "postDate": "2021-06-09T21:31:23.030Z",
          "content": "<p>Not the size of the training data, but the validation. However, the size of that depends on the number of folds, so don't fix the size of the benchmark. (Also, the slowest implementations are as good as the fastest ones, as time spent on calculating auc is so small compared to training… :D)</p>\n<p>But why do you want to use only CPU?</p>",
          "rawMarkdown": "Not the size of the training data, but the validation. However, the size of that depends on the number of folds, so don't fix the size of the benchmark. (Also, the slowest implementations are as good as the fastest ones, as time spent on calculating auc is so small compared to training... :D)\n\nBut why do you want to use only CPU?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1338909,
      "postDate": "2021-06-06T21:05:50.503Z",
      "content": "<p>Perhaps an if-statement could speed things up some more? I.e. if y_i then auc += nfalse else nfalse++ ?</p>",
      "rawMarkdown": "Perhaps an if-statement could speed things up some more? I.e. if y_i then auc += nfalse else nfalse++ ?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1338916,
          "postDate": "2021-06-06T21:13:12.507Z",
          "content": "<p>That's clever!</p>",
          "rawMarkdown": "That's clever!"
        },
        {
          "id": 1338920,
          "postDate": "2021-06-06T21:18:05.933Z",
          "content": "<p>Tried it on 10k elements, but it makes no difference.<br>\nPython works in mysterious ways.</p>",
          "rawMarkdown": "Tried it on 10k elements, but it makes no difference.\nPython works in mysterious ways."
        },
        {
          "id": 1338938,
          "postDate": "2021-06-06T21:48:54.833Z",
          "content": "<p><code>if</code> statements can slow things down if the branch predictor is wrong.  My code (and <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> code) replaces it by  a multiplication, which is always faster.</p>",
          "rawMarkdown": "`if` statements can slow things down if the branch predictor is wrong.  My code (and @nofreewill code) replaces it by  a multiplication, which is always faster.",
          "votes": 3
        },
        {
          "id": 1339536,
          "postDate": "2021-06-07T10:00:07.163Z",
          "content": "<p>A tstl &amp; 2 x br vs subl &amp; mull fooled an old schooler. Proof is in the pudding: 100 000 000 items with if : 1.197 ± 0.004 s and with your multiplication: 0.421 ± 0.003 s -- 2.84 ± 0.03 times faster. Thanks! Time to change my Auc-routine…</p>",
          "rawMarkdown": "A tstl & 2 x br vs subl & mull fooled an old schooler. Proof is in the pudding: 100 000 000 items with if : 1.197 ± 0.004 s and with your multiplication: 0.421 ± 0.003 s -- 2.84 ± 0.03 times faster. Thanks! Time to change my Auc-routine...",
          "votes": 2
        },
        {
          "id": 1339628,
          "postDate": "2021-06-07T11:05:15.257Z",
          "content": "<blockquote>\n  <p>tstl &amp; 2 x br vs subl &amp; mull</p>\n</blockquote>\n<p>It's some form of elvish. I can't read it.</p>",
          "rawMarkdown": "> tstl & 2 x br vs subl & mull\n\nIt's some form of elvish. I can't read it.",
          "votes": 1
        },
        {
          "id": 1339679,
          "postDate": "2021-06-07T11:47:55.750Z",
          "content": "<p>May I ask, what does that mean? ^^'</p>",
          "rawMarkdown": "May I ask, what does that mean? ^^'",
          "votes": 1
        },
        {
          "id": 1339821,
          "postDate": "2021-06-07T13:26:35.903Z",
          "content": "<p>I did a naive machine code of the flow.</p>\n<p><code>tstl yi</code><br>\n<code>brc $1</code><br>\n<code>addl nfalse, auc</code><br>\n<code>br $2</code><br>\n<code>$1:</code><br>\n<code>inc nfalse</code><br>\n<code>$2:</code></p>\n<p>versus</p>\n<p><code>subl yi,#1,yi</code><br>\n<code>addl nfalse,yi</code><br>\n<code>mull yi,nfalse</code><br>\n<code>addl auc,yi</code></p>\n<p>So, the difference would be \"tstl, 2 x br\" vs \"subl, mull\", where I thought a test and two branchings would be faster than a subtraction + multiplication, but not so.</p>",
          "rawMarkdown": "I did a naive machine code of the flow.\n\n`tstl yi`\n`brc $1`\n`addl nfalse, auc`\n`br $2`\n`$1:`\n`inc nfalse`\n`$2:`\n\nversus\n\n`subl yi,#1,yi`\n`addl nfalse,yi`\n`mull yi,nfalse`\n`addl auc,yi`\n\nSo, the difference would be \"tstl, 2 x br\" vs \"subl, mull\", where I thought a test and two branchings would be faster than a subtraction + multiplication, but not so.\n",
          "votes": 3
        },
        {
          "id": 1339896,
          "postDate": "2021-06-07T14:15:52.197Z",
          "content": "<p>Now I see. Did you try it in pure assembly?</p>",
          "rawMarkdown": "Now I see. Did you try it in pure assembly?"
        },
        {
          "id": 1340008,
          "postDate": "2021-06-07T15:34:01.437Z",
          "content": "<p>No, this was just a \"thought experiment\". I implemented it in C# (VS) for AUC ~ 0.5, i.e. random array. I find it interesting to try for AUC ~ 0.9.</p>",
          "rawMarkdown": "No, this was just a \"thought experiment\". I implemented it in C# (VS) for AUC ~ 0.5, i.e. random array. I find it interesting to try for AUC ~ 0.9."
        }
      ]
    },
    {
      "id": 1471119,
      "postDate": "2021-08-14T01:10:15.080Z",
      "content": "<p>When I changed the metric code from sklearn's one to using fast_auc_aux one, val_loss is nan but it goes to the next epoch. It's strange…</p>",
      "rawMarkdown": "When I changed the metric code from sklearn's one to using fast_auc_aux one, val_loss is nan but it goes to the next epoch. It's strange...",
      "replies": [
        {
          "id": 1475732,
          "postDate": "2021-08-16T20:12:11.017Z",
          "content": "<p>fast_auc_score is not a loss, how could it impact your loss?</p>",
          "rawMarkdown": "fast_auc_score is not a loss, how could it impact your loss?"
        },
        {
          "id": 1475828,
          "postDate": "2021-08-16T21:01:41.467Z",
          "content": "<pre><code>[2021-08-11 16:03:52,981][__main__][INFO] -   Epoch  - avg_train_loss: 0.000195  avg_val_loss: nan\n[2021-08-11 16:03:52,982][__main__][INFO] -   Epoch  - AUC : 0.876979\n</code></pre>\n<p>I got an above log when I  use metric=fast_auc_score, criterion=nn.CrossEntropyLoss.</p>\n<p>When I use metric=sklearn.metrics.roc_auc_score, criterion=nn.CrossEntropyLoss.<br>\nI got the error like this if avg_val_loss was nan</p>\n<pre><code>Input contains NaN, infinity or a value too large for dtype('float16')\n</code></pre>",
          "rawMarkdown": "```\n[2021-08-11 16:03:52,981][__main__][INFO] -   Epoch  - avg_train_loss: 0.000195  avg_val_loss: nan\n[2021-08-11 16:03:52,982][__main__][INFO] -   Epoch  - AUC : 0.876979\n```\nI got an above log when I  use metric=fast_auc_score, criterion=nn.CrossEntropyLoss.\n\nWhen I use metric=sklearn.metrics.roc_auc_score, criterion=nn.CrossEntropyLoss.\nI got the error like this if avg_val_loss was nan\n```\nInput contains NaN, infinity or a value too large for dtype('float16')\n```"
        },
        {
          "id": 1476289,
          "postDate": "2021-08-17T04:28:25.327Z",
          "content": "<p>are you saying that what you call avg_val_loss is roc-auc? Name is misleading as roc-auc is not a loss function, it is not differentiable.</p>\n<p>Anyway, the error is clear: you should cast your output to float32 before calling roc-auc functions, be it mine or sklearn one.  fp16 overflows when computing it.</p>",
          "rawMarkdown": "are you saying that what you call avg_val_loss is roc-auc? Name is misleading as roc-auc is not a loss function, it is not differentiable.\n\nAnyway, the error is clear: you should cast your output to float32 before calling roc-auc functions, be it mine or sklearn one.  fp16 overflows when computing it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1376023,
      "postDate": "2021-07-04T17:18:46.613Z",
      "content": "<pre><code>from sklearn.metrics import roc_auc_score\n\n@numba.jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n\n@numba.jit\ndef fast_auc_cpmp(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\n\ny_true = np.array([0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0\n, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0\n, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], dtype=np.int64)\n\ny_prob = np.array([0.5097656,0.45507812,0.5205078,0.50146484,0.49243164,0.5131836\n,0.45117188,0.5546875,0.49682617,0.5107422,0.50683594,0.45361328\n,0.4650879,0.5078125,0.52783203,0.60302734,0.48291016,0.50341797\n,0.5810547,0.45507812,0.48413086,0.5444336,0.5,0.5498047\n,0.5620117,0.47338867,0.47143555,0.58154297,0.5727539,0.5419922\n,0.5571289,0.59033203,0.54589844,0.5463867,0.50146484,0.54785156\n,0.5566406,0.4326172,0.5517578,0.42407227,0.47192383,0.43701172\n,0.5283203,0.46655273,0.5083008,0.5644531,0.55908203,0.515625\n,0.5263672,0.5131836,0.53808594,0.41918945,0.46484375,0.52978516\n,0.5410156,0.41796875,0.57666016,0.4675293,0.5776367,0.52490234\n,0.46606445,0.4819336,0.53125,0.5336914,0.4909668,0.52197266\n,0.49780273,0.55615234,0.5131836,0.51660156,0.43188477,0.57177734\n,0.54296875,0.5683594,0.42529297,0.47143555,0.5678711,0.4753418\n,0.5214844,0.47021484,0.42822266,0.42529297,0.50927734,0.5253906\n,0.51171875,0.46166992,0.59375,0.56884766,0.45629883,0.51708984\n,0.45703125,0.51953125,0.5703125,0.5595703,0.55810547,0.46972656\n,0.4609375,0.5,0.60839844,0.48168945], dtype=np.float32)\n\nprint(roc_auc_score(y_true, y_prob))  # 0.5495923913043478\nprint(fast_auc(y_true, y_prob))       # 0.5516304347826086\nprint(fast_auc_cpmp(y_true, y_prob))  # 0.5489130434782609\n</code></pre>\n<p>Are these results intended? With an even larger number of samples, I've seen the divergence be as much as 0.3 AUC.</p>",
      "rawMarkdown": "```\nfrom sklearn.metrics import roc_auc_score\n\n@numba.jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n\n@numba.jit\ndef fast_auc_cpmp(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\n\ny_true = np.array([0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0\n, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0\n, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], dtype=np.int64)\n\ny_prob = np.array([0.5097656,0.45507812,0.5205078,0.50146484,0.49243164,0.5131836\n,0.45117188,0.5546875,0.49682617,0.5107422,0.50683594,0.45361328\n,0.4650879,0.5078125,0.52783203,0.60302734,0.48291016,0.50341797\n,0.5810547,0.45507812,0.48413086,0.5444336,0.5,0.5498047\n,0.5620117,0.47338867,0.47143555,0.58154297,0.5727539,0.5419922\n,0.5571289,0.59033203,0.54589844,0.5463867,0.50146484,0.54785156\n,0.5566406,0.4326172,0.5517578,0.42407227,0.47192383,0.43701172\n,0.5283203,0.46655273,0.5083008,0.5644531,0.55908203,0.515625\n,0.5263672,0.5131836,0.53808594,0.41918945,0.46484375,0.52978516\n,0.5410156,0.41796875,0.57666016,0.4675293,0.5776367,0.52490234\n,0.46606445,0.4819336,0.53125,0.5336914,0.4909668,0.52197266\n,0.49780273,0.55615234,0.5131836,0.51660156,0.43188477,0.57177734\n,0.54296875,0.5683594,0.42529297,0.47143555,0.5678711,0.4753418\n,0.5214844,0.47021484,0.42822266,0.42529297,0.50927734,0.5253906\n,0.51171875,0.46166992,0.59375,0.56884766,0.45629883,0.51708984\n,0.45703125,0.51953125,0.5703125,0.5595703,0.55810547,0.46972656\n,0.4609375,0.5,0.60839844,0.48168945], dtype=np.float32)\n\nprint(roc_auc_score(y_true, y_prob))  # 0.5495923913043478\nprint(fast_auc(y_true, y_prob))       # 0.5516304347826086\nprint(fast_auc_cpmp(y_true, y_prob))  # 0.5489130434782609\n```\n\nAre these results intended? With an even larger number of samples, I've seen the divergence be as much as 0.3 AUC.",
      "replies": [
        {
          "id": 1376188,
          "postDate": "2021-07-04T22:36:09.207Z",
          "content": "<p>I need to dig into it.</p>",
          "rawMarkdown": "I need to dig into it."
        },
        {
          "id": 1376848,
          "postDate": "2021-07-05T11:41:35.967Z",
          "content": "<p>Run fast_auc_cpmp(.), then fast_auc(.) ?</p>",
          "rawMarkdown": "Run fast_auc_cpmp(.), then fast_auc(.) ?",
          "votes": 1
        },
        {
          "id": 1376857,
          "postDate": "2021-07-05T11:46:43.080Z",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> You have repeated values in y_prob.  The sorting algorithm break ties in a different way, and this changes the score.  Repeated values:</p>\n<pre><code>0.51318359375\n0.471435546875\n0.42529296875\n0.455078125\n0.50146484375\n0.5\n</code></pre>\n<p>I doubt you got 0.3 auc difference.  Isn't it 0.03 rather?</p>\n<p>sklearn deals with repeated probas explicitly.</p>",
          "rawMarkdown": "@authman You have repeated values in y_prob.  The sorting algorithm break ties in a different way, and this changes the score.  Repeated values:\n\n```\n0.51318359375\n0.471435546875\n0.42529296875\n0.455078125\n0.50146484375\n0.5\n```\n\nI doubt you got 0.3 auc difference.  Isn't it 0.03 rather?\n\nsklearn deals with repeated probas explicitly.",
          "votes": 2
        },
        {
          "id": 1378461,
          "postDate": "2021-07-06T14:49:27.137Z",
          "content": "<p>The AUC-error should be around (1/8+1/92)/4~0.034</p>",
          "rawMarkdown": "The AUC-error should be around (1/8+1/92)/4~0.034",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1338610,
      "author_name": "nofreewill42",
      "author_url": "",
      "post_date": "2021-06-06T15:28:19.087000",
      "content": "<p>Speeking of speed, there is also np.cumsum:</p>\n<pre><code>%%timeit\nnp.cumsum(np.arange(1000))\n</code></pre>\n<p>4.01 µs ± 14.4 ns</p>\n<pre><code>%%timeit\ns = 0\nfor i in range(1000):\n    s += i\n</code></pre>\n<p>35.4 µs ± 229 ns (EDIT: note that numba's jit decorator speeds up the original loop quite a bit!)</p>\n<p>Code utilizing it:</p>\n<pre><code>def fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>",
      "votes": 9,
      "replies": [
        {
          "id": 1338877,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-06T20:05:39.077000",
          "content": "<p>Your fast auc is 16% faster indeed.  Thanks for sharing.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1338886,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-06T20:20:13.153000",
          "content": "<p>On an array of 10000 elements:<br>\n[cumsum vs for loop]</p>\n<p><strong>Without jit:</strong><br>\n[<strong>575 µs</strong> vs 3.36 ms]</p>\n<p><strong>With jit:</strong><br>\n[727 µs vs 709 µs]</p>\n<p>So the best time is <strong>cumsum without numba's jit</strong> decorator, although not at all that much as those initial timings would suggest.<br>\nI didn't know numba, so thank you for introducing it. It smells like magic :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1338892,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-06T20:37:42.987000",
          "content": "<p>On an array of 1000 elements, though,<br>\nyour for loop with jit is a little better than cumsum. Both are around 37 us.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1338936,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-06T21:47:29.843000",
          "content": "<p>yeah, I found the same.  cumsum wins as the array gets larger.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339681,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-07T11:49:05.323000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> I made numba code run faster than yours (by a small margin):</p>\n<pre><code>@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339682,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T11:52:31.550000",
          "content": "<p>Challange accepted! :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339780,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T12:59:34.417000",
          "content": "<p>I have no gpu to try it right now, but I'm hopeful that it's faster.<br>\nTill then, here is the code.</p>\n<pre><code>def fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n</code></pre>\n<p>One should collect their y_probs and y_trues on the gpu as they predict, right away, to bypass additional cpu -&gt; gpu communications.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1340053,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T16:19:28.657000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I could make it faster on gpu with torch's cumsum</p>\n<p><strong>data</strong></p>\n<pre><code>y_true = np.random.randint(2,size=10000)\ny_prob = np.random.rand(10000)\n</code></pre>\n<p><strong>your most recent version</strong></p>\n<pre><code>@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n</code></pre>\n<pre><code>%%timeit\nfast_auc(y_true, y_prob)\n</code></pre>\n<p>537 µs</p>\n<p><strong>fast_auc_torch</strong> - sending to gpu first</p>\n<pre><code>%%timeit\nfast_auc_torch(torch.tensor(y_true).to(device), torch.tensor(y_prob).to(device))\n</code></pre>\n<p>340 µs</p>\n<p><strong>fast_auc_torch</strong> - data is already on gpu</p>\n<pre><code>y_true_tensor = torch.tensor(y_true).to(device), \ny_prob_tensor = torch.tensor(y_prob).to(device)\n</code></pre>\n<pre><code>%%timeit\nfast_auc_torch(y_true_tensor, y_prob_tensor)\n</code></pre>\n<p>205 µs</p>\n<p>On a 1080Ti</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1340058,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T16:25:06.827000",
          "content": "<p>With 1M number of elements.<br>\nsklearn: <strong>261 ms</strong><br>\nnumpy cumsum: <strong>100 ms</strong><br>\nfast_auc_forloop with jit: <strong>131 ms</strong><br>\nfast_auc with separate fast_auc_aux: <strong>95.2 ms</strong><br>\nfast_auc_torch sending data to gpu: <strong>10 ms</strong><br>\nfast_auc_torch data is already on gpu: <strong>2.83 ms</strong></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1340063,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T16:32:07.820000",
          "content": "<p>sklearn on 10k takes 2.4 ms, so in this competition with a few 10k samples, we can save a second in every couple of hundred epochs … xd</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1341504,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2021-06-08T17:23:46.510000",
          "content": "<blockquote>\n  <p>Challange accepted! :D</p>\n</blockquote>\n<p>Let's show the Aliens how fast we can roc-auc them ;)</p>\n<p>P.S. I suggest to use CPU only and seti train dataset shape for benchmarking our implementations. Are you?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1342947,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-09T21:31:23.030000",
          "content": "<p>Not the size of the training data, but the validation. However, the size of that depends on the number of folds, so don't fix the size of the benchmark. (Also, the slowest implementations are as good as the fastest ones, as time spent on calculating auc is so small compared to training… :D)</p>\n<p>But why do you want to use only CPU?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1338909,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2021-06-06T21:05:50.503000",
      "content": "<p>Perhaps an if-statement could speed things up some more? I.e. if y_i then auc += nfalse else nfalse++ ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1338916,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-06T21:13:12.507000",
          "content": "<p>That's clever!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1338920,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-06T21:18:05.933000",
          "content": "<p>Tried it on 10k elements, but it makes no difference.<br>\nPython works in mysterious ways.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1338938,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-06-06T21:48:54.833000",
          "content": "<p><code>if</code> statements can slow things down if the branch predictor is wrong.  My code (and <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> code) replaces it by  a multiplication, which is always faster.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1339536,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-06-07T10:00:07.163000",
          "content": "<p>A tstl &amp; 2 x br vs subl &amp; mull fooled an old schooler. Proof is in the pudding: 100 000 000 items with if : 1.197 ± 0.004 s and with your multiplication: 0.421 ± 0.003 s -- 2.84 ± 0.03 times faster. Thanks! Time to change my Auc-routine…</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1339628,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T11:05:15.257000",
          "content": "<blockquote>\n  <p>tstl &amp; 2 x br vs subl &amp; mull</p>\n</blockquote>\n<p>It's some form of elvish. I can't read it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339679,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T11:47:55.750000",
          "content": "<p>May I ask, what does that mean? ^^'</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1339821,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-06-07T13:26:35.903000",
          "content": "<p>I did a naive machine code of the flow.</p>\n<p><code>tstl yi</code><br>\n<code>brc $1</code><br>\n<code>addl nfalse, auc</code><br>\n<code>br $2</code><br>\n<code>$1:</code><br>\n<code>inc nfalse</code><br>\n<code>$2:</code></p>\n<p>versus</p>\n<p><code>subl yi,#1,yi</code><br>\n<code>addl nfalse,yi</code><br>\n<code>mull yi,nfalse</code><br>\n<code>addl auc,yi</code></p>\n<p>So, the difference would be \"tstl, 2 x br\" vs \"subl, mull\", where I thought a test and two branchings would be faster than a subtraction + multiplication, but not so.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1339896,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-07T14:15:52.197000",
          "content": "<p>Now I see. Did you try it in pure assembly?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1340008,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-06-07T15:34:01.437000",
          "content": "<p>No, this was just a \"thought experiment\". I implemented it in C# (VS) for AUC ~ 0.5, i.e. random array. I find it interesting to try for AUC ~ 0.9.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1471119,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-08-14T01:10:15.080000",
      "content": "<p>When I changed the metric code from sklearn's one to using fast_auc_aux one, val_loss is nan but it goes to the next epoch. It's strange…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1475732,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-16T20:12:11.017000",
          "content": "<p>fast_auc_score is not a loss, how could it impact your loss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1475828,
          "author_name": "patriot",
          "author_url": "",
          "post_date": "2021-08-16T21:01:41.467000",
          "content": "<pre><code>[2021-08-11 16:03:52,981][__main__][INFO] -   Epoch  - avg_train_loss: 0.000195  avg_val_loss: nan\n[2021-08-11 16:03:52,982][__main__][INFO] -   Epoch  - AUC : 0.876979\n</code></pre>\n<p>I got an above log when I  use metric=fast_auc_score, criterion=nn.CrossEntropyLoss.</p>\n<p>When I use metric=sklearn.metrics.roc_auc_score, criterion=nn.CrossEntropyLoss.<br>\nI got the error like this if avg_val_loss was nan</p>\n<pre><code>Input contains NaN, infinity or a value too large for dtype('float16')\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1476289,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-08-17T04:28:25.327000",
          "content": "<p>are you saying that what you call avg_val_loss is roc-auc? Name is misleading as roc-auc is not a loss function, it is not differentiable.</p>\n<p>Anyway, the error is clear: you should cast your output to float32 before calling roc-auc functions, be it mine or sklearn one.  fp16 overflows when computing it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1376023,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2021-07-04T17:18:46.613000",
      "content": "<pre><code>from sklearn.metrics import roc_auc_score\n\n@numba.jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n\n@numba.jit\ndef fast_auc_cpmp(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\n\ny_true = np.array([0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0\n, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0\n, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], dtype=np.int64)\n\ny_prob = np.array([0.5097656,0.45507812,0.5205078,0.50146484,0.49243164,0.5131836\n,0.45117188,0.5546875,0.49682617,0.5107422,0.50683594,0.45361328\n,0.4650879,0.5078125,0.52783203,0.60302734,0.48291016,0.50341797\n,0.5810547,0.45507812,0.48413086,0.5444336,0.5,0.5498047\n,0.5620117,0.47338867,0.47143555,0.58154297,0.5727539,0.5419922\n,0.5571289,0.59033203,0.54589844,0.5463867,0.50146484,0.54785156\n,0.5566406,0.4326172,0.5517578,0.42407227,0.47192383,0.43701172\n,0.5283203,0.46655273,0.5083008,0.5644531,0.55908203,0.515625\n,0.5263672,0.5131836,0.53808594,0.41918945,0.46484375,0.52978516\n,0.5410156,0.41796875,0.57666016,0.4675293,0.5776367,0.52490234\n,0.46606445,0.4819336,0.53125,0.5336914,0.4909668,0.52197266\n,0.49780273,0.55615234,0.5131836,0.51660156,0.43188477,0.57177734\n,0.54296875,0.5683594,0.42529297,0.47143555,0.5678711,0.4753418\n,0.5214844,0.47021484,0.42822266,0.42529297,0.50927734,0.5253906\n,0.51171875,0.46166992,0.59375,0.56884766,0.45629883,0.51708984\n,0.45703125,0.51953125,0.5703125,0.5595703,0.55810547,0.46972656\n,0.4609375,0.5,0.60839844,0.48168945], dtype=np.float32)\n\nprint(roc_auc_score(y_true, y_prob))  # 0.5495923913043478\nprint(fast_auc(y_true, y_prob))       # 0.5516304347826086\nprint(fast_auc_cpmp(y_true, y_prob))  # 0.5489130434782609\n</code></pre>\n<p>Are these results intended? With an even larger number of samples, I've seen the divergence be as much as 0.3 AUC.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1376188,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-04T22:36:09.207000",
          "content": "<p>I need to dig into it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1376848,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-07-05T11:41:35.967000",
          "content": "<p>Run fast_auc_cpmp(.), then fast_auc(.) ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1376857,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2021-07-05T11:46:43.080000",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> You have repeated values in y_prob.  The sorting algorithm break ties in a different way, and this changes the score.  Repeated values:</p>\n<pre><code>0.51318359375\n0.471435546875\n0.42529296875\n0.455078125\n0.50146484375\n0.5\n</code></pre>\n<p>I doubt you got 0.3 auc difference.  Isn't it 0.03 rather?</p>\n<p>sklearn deals with repeated probas explicitly.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1378461,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2021-07-06T14:49:27.137000",
          "content": "<p>The AUC-error should be around (1/8+1/92)/4~0.034</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1338552": "Here is a code that computes roc-auc.  It is way faster than scikit-learn implementation.  However, if you have ties in your predictions then the value might differ from that of scikit-learn.\n\nMake sure you respect the order of parameters: the first one must be the target, the second one the predictions.  You can use logits, there is no need to input probability as predictions.  Indeed, what matters is the ordering of your predictions, not their absolute values.\n\n```\nimport numpy as np \nfrom numba import jit\n\n@jit\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n```\n\nEdt.  @nofreewill improved over it using numpy cumsum, see comments.  I include his code here for convenience.\n\n```\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n\n```\n\nEdit2.  This code is even faster:\n\n```\n@jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n```\n\nEdit 3: @nofreewill has shared that using GPU is even faster.  Here is his code:|\n\n```\n\ndef fast_auc_torch(y_true:torch.Tensor, y_prob:torch.Tensor):\n    y_true = y_true[torch.argsort(y_prob)]\n    cumfalses = torch.cumsum(1-y_true,dim=0)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc = auc / (nfalse * (len(y_true) - nfalse))\n    return auc\n\n```\n",
    "1338610": "Speeking of speed, there is also np.cumsum:\n\n```\n%%timeit\nnp.cumsum(np.arange(1000))\n```\n4.01 µs ± 14.4 ns\n\n```\n%%timeit\ns = 0\nfor i in range(1000):\n    s += i\n```\n35.4 µs ± 229 ns (EDIT: note that numba's jit decorator speeds up the original loop quite a bit!)\n\nCode utilizing it:\n```\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    cumfalses = np.cumsum(1-y_true)\n    nfalse = cumfalses[-1]\n    auc = (y_true * cumfalses).sum()\n    auc /= (nfalse * (len(y_true) - nfalse))\n    return auc\n```",
    "1338909": "Perhaps an if-statement could speed things up some more? I.e. if y_i then auc += nfalse else nfalse++ ?\n",
    "1471119": "When I changed the metric code from sklearn's one to using fast_auc_aux one, val_loss is nan but it goes to the next epoch. It's strange...",
    "1376023": "```\nfrom sklearn.metrics import roc_auc_score\n\n@numba.jit()\ndef fast_auc_aux(y_true, y_prob):\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\ndef fast_auc(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    return fast_auc_aux(y_true, y_prob)\n\n@numba.jit\ndef fast_auc_cpmp(y_true, y_prob):\n    y_true = np.asarray(y_true)\n    y_true = y_true[np.argsort(y_prob)]\n    nfalse = 0\n    auc = 0\n    n = len(y_true)\n    for i in range(n):\n        y_i = y_true[i]\n        nfalse += (1 - y_i)\n        auc += y_i * nfalse\n    auc /= (nfalse * (n - nfalse))\n    return auc\n\n\ny_true = np.array([0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0\n, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 0, 0, 0, 0\n, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0], dtype=np.int64)\n\ny_prob = np.array([0.5097656,0.45507812,0.5205078,0.50146484,0.49243164,0.5131836\n,0.45117188,0.5546875,0.49682617,0.5107422,0.50683594,0.45361328\n,0.4650879,0.5078125,0.52783203,0.60302734,0.48291016,0.50341797\n,0.5810547,0.45507812,0.48413086,0.5444336,0.5,0.5498047\n,0.5620117,0.47338867,0.47143555,0.58154297,0.5727539,0.5419922\n,0.5571289,0.59033203,0.54589844,0.5463867,0.50146484,0.54785156\n,0.5566406,0.4326172,0.5517578,0.42407227,0.47192383,0.43701172\n,0.5283203,0.46655273,0.5083008,0.5644531,0.55908203,0.515625\n,0.5263672,0.5131836,0.53808594,0.41918945,0.46484375,0.52978516\n,0.5410156,0.41796875,0.57666016,0.4675293,0.5776367,0.52490234\n,0.46606445,0.4819336,0.53125,0.5336914,0.4909668,0.52197266\n,0.49780273,0.55615234,0.5131836,0.51660156,0.43188477,0.57177734\n,0.54296875,0.5683594,0.42529297,0.47143555,0.5678711,0.4753418\n,0.5214844,0.47021484,0.42822266,0.42529297,0.50927734,0.5253906\n,0.51171875,0.46166992,0.59375,0.56884766,0.45629883,0.51708984\n,0.45703125,0.51953125,0.5703125,0.5595703,0.55810547,0.46972656\n,0.4609375,0.5,0.60839844,0.48168945], dtype=np.float32)\n\nprint(roc_auc_score(y_true, y_prob))  # 0.5495923913043478\nprint(fast_auc(y_true, y_prob))       # 0.5516304347826086\nprint(fast_auc_cpmp(y_true, y_prob))  # 0.5489130434782609\n```\n\nAre these results intended? With an even larger number of samples, I've seen the divergence be as much as 0.3 AUC."
  }
}