{
  "id": 328020,
  "title": "10x fast metric (numpy)",
  "url": "/competitions/amex-default-prediction/discussion/328020",
  "author_name": "Pims",
  "post_date": "2022-05-30T14:20:58.852000",
  "votes": 111,
  "comment_count": 13,
  "views": 0,
  "content": "<pre><code>def amex_metric_np(preds: np.ndarray, target: np.ndarray) -&gt; float:\n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n\n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight / weight.sum()).cumsum()\n    four_pct_mask = cum_norm_weight &lt;= 0.04\n    d = np.sum(target[four_pct_mask]) / np.sum(target)\n\n    weighted_target = target * weight\n    lorentz = (weighted_target / weighted_target.sum()).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n    gini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)\n\n    g = gini / gini_max\n    return 0.5 * (g + d)\n</code></pre>",
  "messages": [
    {
      "id": 1805812,
      "postDate": "2022-05-30T14:20:58.853Z",
      "content": "<pre><code>def amex_metric_np(preds: np.ndarray, target: np.ndarray) -&gt; float:\n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n\n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight / weight.sum()).cumsum()\n    four_pct_mask = cum_norm_weight &lt;= 0.04\n    d = np.sum(target[four_pct_mask]) / np.sum(target)\n\n    weighted_target = target * weight\n    lorentz = (weighted_target / weighted_target.sum()).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n    gini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)\n\n    g = gini / gini_max\n    return 0.5 * (g + d)\n</code></pre>",
      "rawMarkdown": "```python\ndef amex_metric_np(preds: np.ndarray, target: np.ndarray) -> float:\n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n    \n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight / weight.sum()).cumsum()\n    four_pct_mask = cum_norm_weight <= 0.04\n    d = np.sum(target[four_pct_mask]) / np.sum(target)\n    \n    weighted_target = target * weight\n    lorentz = (weighted_target / weighted_target.sum()).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n    \n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n    gini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)\n    \n    g = gini / gini_max\n    return 0.5 * (g + d)\n```",
      "votes": 111
    },
    {
      "id": 1806309,
      "postDate": "2022-05-31T03:36:23.903Z",
      "content": "<pre><code>def amex_metric_np(preds: np.ndarray, target: np.ndarray) -&gt; float:\n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n\n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n\n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight * (1 / weight.sum())).cumsum()\n    four_pct_mask = cum_norm_weight &lt;= 0.04\n    d = np.sum(target[four_pct_mask]) / n_pos\n\n    lorentz = (target * (1 / n_pos)).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    gini_max = 10 * n_neg * (1 - 19 / (n_pos + 20 * n_neg))\n\n    g = gini / gini_max\n    return 0.5 * (g + d)\n</code></pre>\n<p>can be simplified and faster</p>",
      "rawMarkdown": "```python\ndef amex_metric_np(preds: np.ndarray, target: np.ndarray) -> float:\n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n    \n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n    \n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight * (1 / weight.sum())).cumsum()\n    four_pct_mask = cum_norm_weight <= 0.04\n    d = np.sum(target[four_pct_mask]) / n_pos\n    \n    lorentz = (target * (1 / n_pos)).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    gini_max = 10 * n_neg * (1 - 19 / (n_pos + 20 * n_neg))\n    \n    g = gini / gini_max\n    return 0.5 * (g + d)\n```\n\ncan be simplified and faster",
      "votes": 4
    },
    {
      "id": 1805816,
      "postDate": "2022-05-30T14:24:39.350Z",
      "content": "<p>If you find it helpful please upvote~</p>",
      "rawMarkdown": "If you find it helpful please upvote~",
      "votes": 1
    },
    {
      "id": 1845267,
      "postDate": "2022-07-06T07:20:11Z",
      "content": "<p>Note, other implementations of metrics, including sklearn's, have the ground truth <code>target</code> as the first argument.</p>",
      "rawMarkdown": "Note, other implementations of metrics, including sklearn's, have the ground truth `target` as the first argument."
    },
    {
      "id": 1808040,
      "postDate": "2022-06-01T13:49:24.107Z",
      "content": "<p>Thanks a lot for sharing this!</p>\n<p>The <code>gini_max</code> computation is very cool and brings most of the speed difference.</p>\n<p>I added it to my <a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\">notebook here</a> and the benchmark section shows it's almost 20x faster than the official pandas implementation.</p>",
      "rawMarkdown": "Thanks a lot for sharing this!\n\nThe `gini_max` computation is very cool and brings most of the speed difference.\n\nI added it to my [notebook here](https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations) and the benchmark section shows it's almost 20x faster than the official pandas implementation.",
      "votes": 2
    },
    {
      "id": 1959631,
      "postDate": "2022-09-28T08:20:05.890Z",
      "content": "<p>Hi, Pims, could you please kindly explain the caoculation of gini_max?<br>\nthanks sooooo much<br>\n真的不知道咋算的求解答</p>",
      "rawMarkdown": "Hi, Pims, could you please kindly explain the caoculation of gini_max?\nthanks sooooo much\n真的不知道咋算的求解答"
    },
    {
      "id": 1904528,
      "postDate": "2022-08-18T08:58:03.183Z",
      "content": "<p>Hi, Pims, could you explain the calculation of gini_max? Kinda of curious how to get this:</p>\n<p>gini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)</p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Hi, Pims, could you explain the calculation of gini_max? Kinda of curious how to get this:\n\ngini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)\n\nThank you in advance!"
    },
    {
      "id": 1807957,
      "postDate": "2022-06-01T13:00:03.887Z",
      "content": "<p>That's a formula for an evaluation function using Numpy!<br>\nInteresting. Thanks for sharing!👍</p>",
      "rawMarkdown": "That's a formula for an evaluation function using Numpy!\nInteresting. Thanks for sharing!👍"
    },
    {
      "id": 1883587,
      "postDate": "2022-08-03T23:33:42.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1883712,
      "postDate": "2022-08-04T03:33:50.343Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 1859365,
      "postDate": "2022-07-17T15:09:45.373Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 1850780,
      "postDate": "2022-07-10T18:27:45.130Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!"
    },
    {
      "id": 1806631,
      "postDate": "2022-05-31T10:50:50.920Z",
      "content": "<p>thanks for the info.👍</p>",
      "rawMarkdown": "thanks for the info.👍"
    },
    {
      "id": 1806262,
      "postDate": "2022-05-31T01:25:39.413Z",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing! "
    }
  ],
  "comments": [
    {
      "id": 1806309,
      "author_name": "Pims",
      "author_url": "",
      "post_date": "2022-05-31T03:36:23.903000",
      "content": "<pre><code>def amex_metric_np(preds: np.ndarray, target: np.ndarray) -&gt; float:\n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n\n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n\n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight * (1 / weight.sum())).cumsum()\n    four_pct_mask = cum_norm_weight &lt;= 0.04\n    d = np.sum(target[four_pct_mask]) / n_pos\n\n    lorentz = (target * (1 / n_pos)).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    gini_max = 10 * n_neg * (1 - 19 / (n_pos + 20 * n_neg))\n\n    g = gini / gini_max\n    return 0.5 * (g + d)\n</code></pre>\n<p>can be simplified and faster</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1805816,
      "author_name": "Pims",
      "author_url": "",
      "post_date": "2022-05-30T14:24:39.350000",
      "content": "<p>If you find it helpful please upvote~</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1845267,
      "author_name": "Burrito Dan",
      "author_url": "",
      "post_date": "2022-07-06T07:20:11",
      "content": "<p>Note, other implementations of metrics, including sklearn's, have the ground truth <code>target</code> as the first argument.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1808040,
      "author_name": "Vopani",
      "author_url": "",
      "post_date": "2022-06-01T13:49:24.107000",
      "content": "<p>Thanks a lot for sharing this!</p>\n<p>The <code>gini_max</code> computation is very cool and brings most of the speed difference.</p>\n<p>I added it to my <a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\">notebook here</a> and the benchmark section shows it's almost 20x faster than the official pandas implementation.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1959631,
      "author_name": "YuxuanDeng",
      "author_url": "",
      "post_date": "2022-09-28T08:20:05.890000",
      "content": "<p>Hi, Pims, could you please kindly explain the caoculation of gini_max?<br>\nthanks sooooo much<br>\n真的不知道咋算的求解答</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1904528,
      "author_name": "dehiker",
      "author_url": "",
      "post_date": "2022-08-18T08:58:03.183000",
      "content": "<p>Hi, Pims, could you explain the calculation of gini_max? Kinda of curious how to get this:</p>\n<p>gini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)</p>\n<p>Thank you in advance!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1807957,
      "author_name": "chinchilla",
      "author_url": "",
      "post_date": "2022-06-01T13:00:03.887000",
      "content": "<p>That's a formula for an evaluation function using Numpy!<br>\nInteresting. Thanks for sharing!👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1883587,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-03T23:33:42.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1883712,
      "author_name": "Pengfei",
      "author_url": "",
      "post_date": "2022-08-04T03:33:50.343000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1859365,
      "author_name": "Rafail Giavrimis",
      "author_url": "",
      "post_date": "2022-07-17T15:09:45.373000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1850780,
      "author_name": "Toprak Ucar",
      "author_url": "",
      "post_date": "2022-07-10T18:27:45.130000",
      "content": "<p>thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1806631,
      "author_name": "Suraj Krishna_123",
      "author_url": "",
      "post_date": "2022-05-31T10:50:50.920000",
      "content": "<p>thanks for the info.👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1806262,
      "author_name": "Tonghui Li",
      "author_url": "",
      "post_date": "2022-05-31T01:25:39.413000",
      "content": "<p>Thanks for sharing! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1805812": "```python\ndef amex_metric_np(preds: np.ndarray, target: np.ndarray) -> float:\n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n    \n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight / weight.sum()).cumsum()\n    four_pct_mask = cum_norm_weight <= 0.04\n    d = np.sum(target[four_pct_mask]) / np.sum(target)\n    \n    weighted_target = target * weight\n    lorentz = (weighted_target / weighted_target.sum()).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n    \n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n    gini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)\n    \n    g = gini / gini_max\n    return 0.5 * (g + d)\n```",
    "1806309": "```python\ndef amex_metric_np(preds: np.ndarray, target: np.ndarray) -> float:\n    n_pos = np.sum(target)\n    n_neg = target.shape[0] - n_pos\n    \n    indices = np.argsort(preds)[::-1]\n    preds, target = preds[indices], target[indices]\n    \n    weight = 20.0 - target * 19.0\n    cum_norm_weight = (weight * (1 / weight.sum())).cumsum()\n    four_pct_mask = cum_norm_weight <= 0.04\n    d = np.sum(target[four_pct_mask]) / n_pos\n    \n    lorentz = (target * (1 / n_pos)).cumsum()\n    gini = ((lorentz - cum_norm_weight) * weight).sum()\n\n    gini_max = 10 * n_neg * (1 - 19 / (n_pos + 20 * n_neg))\n    \n    g = gini / gini_max\n    return 0.5 * (g + d)\n```\n\ncan be simplified and faster",
    "1805816": "If you find it helpful please upvote~",
    "1845267": "Note, other implementations of metrics, including sklearn's, have the ground truth `target` as the first argument.",
    "1808040": "Thanks a lot for sharing this!\n\nThe `gini_max` computation is very cool and brings most of the speed difference.\n\nI added it to my [notebook here](https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations) and the benchmark section shows it's almost 20x faster than the official pandas implementation.",
    "1959631": "Hi, Pims, could you please kindly explain the caoculation of gini_max?\nthanks sooooo much\n真的不知道咋算的求解答",
    "1904528": "Hi, Pims, could you explain the calculation of gini_max? Kinda of curious how to get this:\n\ngini_max = 10 * n_neg * (n_pos + 20 * n_neg - 19) / (n_pos + 20 * n_neg)\n\nThank you in advance!",
    "1807957": "That's a formula for an evaluation function using Numpy!\nInteresting. Thanks for sharing!👍",
    "1883587": "",
    "1883712": "Thanks for sharing!",
    "1859365": "Thanks for sharing.",
    "1850780": "thanks for sharing!",
    "1806631": "thanks for the info.👍",
    "1806262": "Thanks for sharing! "
  }
}