{
  "id": 327609,
  "title": "Custom metric without pandas DataFrame, and accelerated by numba",
  "url": "/competitions/amex-default-prediction/discussion/327609",
  "author_name": "",
  "post_date": "2022-05-28T04:31:19.893043400Z",
  "votes": 9,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Since the original AMEX metric is achived using pandas, it is limited when you use is as an early stopping metric in XGBoost or Keras.</p>\n<p>I rewrite the metric using the numpy, and tested it with some unit test cases. Furthermore, the metric is accelerated by numba, which can make the computation much faster.</p>\n<p><a href=\"https://www.kaggle.com/njit\" target=\"_blank\">@njit</a><br>\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):</p>\n<pre><code>pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\ny_true_label_sorted = y_true_label[pred_sorted_idx]\n\nweights = np.ones((len(y_true_label_sorted), ))\nfor i in range(len(weights)):\n    if y_true_label_sorted[i] == 0:\n        weights[i] = 20\n\n# compute cut-off point\nweights_cumsum = np.cumsum(weights)\nfour_pct_cutoff = int(0.04 * np.sum(weights))\n\nfor i in range(len(weights)):\n    if weights_cumsum[i] &gt; four_pct_cutoff:\n        break\ny_true_label_sorted_cut = y_true_label_sorted[:i]\n\n# compute final score\ncutoff_res = np.sum(y_true_label_sorted_cut == 1)\nreal_res = np.sum(y_true_label_sorted == 1)\n\nif real_res == 0:\n    return 0\nelse:\n    return cutoff_res / real_res\n</code></pre>\n<p><a href=\"https://www.kaggle.com/njit\" target=\"_blank\">@njit</a><br>\ndef compute_weighted_gini(y_true_label, y_pred_proba):</p>\n<pre><code>pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\ny_true_label_sorted = y_true_label[pred_sorted_idx]\n\nweights = np.ones((len(y_true_label_sorted), ))\nfor i in range(len(weights)):\n    if y_true_label_sorted[i] == 0:\n        weights[i] = 20\n\nrandoms = np.cumsum(weights / np.sum(weights))\ntotal_pos = np.sum(y_true_label_sorted * weights)\ncum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n# compute final score\nlorentz = cum_pos_found / total_pos\ngini = (lorentz - randoms) * weights\n\nreturn np.sum(gini)\n</code></pre>\n<p>def custom_metric(y_true_label, y_pred_proba):</p>\n<pre><code># compute normalized_weighted_gini score\ntmp_0 = compute_weighted_gini(\n    y_true_label, y_pred_proba\n)\ntmp_1 = compute_weighted_gini(\n    y_true_label, y_true_label\n)\n\nif tmp_1 == 0:\n    g = 0\nelse:\n    g = tmp_0 / tmp_1\n\n# compute top_four_percent_captured score\nd = compute_top_four_percent_captured(\n    y_true_label, y_pred_proba\n)\n\nreturn 0.5 * (g + d)\n</code></pre>",
  "messages": [
    {
      "id": "1803678",
      "postDate": "05/28/2022 04:31:19",
      "content": "<p>Since the original AMEX metric is achived using pandas, it is limited when you use is as an early stopping metric in XGBoost or Keras.</p>\n<p>I rewrite the metric using the numpy, and tested it with some unit test cases. Furthermore, the metric is accelerated by numba, which can make the computation much faster.</p>\n<p><a href=\"https://www.kaggle.com/njit\" target=\"_blank\">@njit</a><br>\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):</p>\n<pre><code>pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\ny_true_label_sorted = y_true_label[pred_sorted_idx]\n\nweights = np.ones((len(y_true_label_sorted), ))\nfor i in range(len(weights)):\n    if y_true_label_sorted[i] == 0:\n        weights[i] = 20\n\n# compute cut-off point\nweights_cumsum = np.cumsum(weights)\nfour_pct_cutoff = int(0.04 * np.sum(weights))\n\nfor i in range(len(weights)):\n    if weights_cumsum[i] &gt; four_pct_cutoff:\n        break\ny_true_label_sorted_cut = y_true_label_sorted[:i]\n\n# compute final score\ncutoff_res = np.sum(y_true_label_sorted_cut == 1)\nreal_res = np.sum(y_true_label_sorted == 1)\n\nif real_res == 0:\n    return 0\nelse:\n    return cutoff_res / real_res\n</code></pre>\n<p><a href=\"https://www.kaggle.com/njit\" target=\"_blank\">@njit</a><br>\ndef compute_weighted_gini(y_true_label, y_pred_proba):</p>\n<pre><code>pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\ny_true_label_sorted = y_true_label[pred_sorted_idx]\n\nweights = np.ones((len(y_true_label_sorted), ))\nfor i in range(len(weights)):\n    if y_true_label_sorted[i] == 0:\n        weights[i] = 20\n\nrandoms = np.cumsum(weights / np.sum(weights))\ntotal_pos = np.sum(y_true_label_sorted * weights)\ncum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n# compute final score\nlorentz = cum_pos_found / total_pos\ngini = (lorentz - randoms) * weights\n\nreturn np.sum(gini)\n</code></pre>\n<p>def custom_metric(y_true_label, y_pred_proba):</p>\n<pre><code># compute normalized_weighted_gini score\ntmp_0 = compute_weighted_gini(\n    y_true_label, y_pred_proba\n)\ntmp_1 = compute_weighted_gini(\n    y_true_label, y_true_label\n)\n\nif tmp_1 == 0:\n    g = 0\nelse:\n    g = tmp_0 / tmp_1\n\n# compute top_four_percent_captured score\nd = compute_top_four_percent_captured(\n    y_true_label, y_pred_proba\n)\n\nreturn 0.5 * (g + d)\n</code></pre>",
      "rawMarkdown": "Since the original AMEX metric is achived using pandas, it is limited when you use is as an early stopping metric in XGBoost or Keras.\n\nI rewrite the metric using the numpy, and tested it with some unit test cases. Furthermore, the metric is accelerated by numba, which can make the computation much faster.\n\n@njit\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):\n\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    # compute cut-off point\n    weights_cumsum = np.cumsum(weights)\n    four_pct_cutoff = int(0.04 * np.sum(weights))\n\n    for i in range(len(weights)):\n        if weights_cumsum[i] > four_pct_cutoff:\n            break\n    y_true_label_sorted_cut = y_true_label_sorted[:i]\n\n    # compute final score\n    cutoff_res = np.sum(y_true_label_sorted_cut == 1)\n    real_res = np.sum(y_true_label_sorted == 1)\n\n    if real_res == 0:\n        return 0\n    else:\n        return cutoff_res / real_res\n\n\n@njit\ndef compute_weighted_gini(y_true_label, y_pred_proba):\n\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    randoms = np.cumsum(weights / np.sum(weights))\n    total_pos = np.sum(y_true_label_sorted * weights)\n    cum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n    # compute final score\n    lorentz = cum_pos_found / total_pos\n    gini = (lorentz - randoms) * weights\n\n    return np.sum(gini)\n\n\ndef custom_metric(y_true_label, y_pred_proba):\n\n    # compute normalized_weighted_gini score\n    tmp_0 = compute_weighted_gini(\n        y_true_label, y_pred_proba\n    )\n    tmp_1 = compute_weighted_gini(\n        y_true_label, y_true_label\n    )\n\n    if tmp_1 == 0:\n        g = 0\n    else:\n        g = tmp_0 / tmp_1\n\n    # compute top_four_percent_captured score\n    d = compute_top_four_percent_captured(\n        y_true_label, y_pred_proba\n    )\n\n    return 0.5 * (g + d)",
      "votes": null
    },
    {
      "id": "1803680",
      "postDate": "05/28/2022 04:32:48",
      "content": "<p>A simple unit test case without assertion:</p>\n<p>def test_custom_metrics():</p>\n<pre><code># test 1\npred_df = pd.DataFrame(\n    {'prediction': np.random.random(2000)}\n)\ntrue_df = pd.DataFrame(\n    {'target': np.random.randint(0, 2, 2000)}\n)\n\nprint(amex_metric(true_df, pred_df))\n\nprint(\n    custom_metric(\n        true_df['target'].values, pred_df['prediction'].values\n    )\n)\n</code></pre>",
      "rawMarkdown": "A simple unit test case without assertion:\n\ndef test_custom_metrics():\n\n    # test 1\n    pred_df = pd.DataFrame(\n        {'prediction': np.random.random(2000)}\n    )\n    true_df = pd.DataFrame(\n        {'target': np.random.randint(0, 2, 2000)}\n    )\n\n    print(amex_metric(true_df, pred_df))\n\n    print(\n        custom_metric(\n            true_df['target'].values, pred_df['prediction'].values\n        )\n    )",
      "votes": null
    },
    {
      "id": "1804318",
      "postDate": "05/28/2022 18:59:10",
      "content": "<p>Slight formatting issue with your discussion post, I believe you meant this:</p>\n<pre><code>@njit()\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    # compute cut-off point\n    weights_cumsum = np.cumsum(weights)\n    four_pct_cutoff = int(0.04 * np.sum(weights))\n\n    for i in range(len(weights)):\n        if weights_cumsum[i] &gt; four_pct_cutoff:\n            break\n    y_true_label_sorted_cut = y_true_label_sorted[:i]\n\n    # compute final score\n    cutoff_res = np.sum(y_true_label_sorted_cut == 1)\n    real_res = np.sum(y_true_label_sorted == 1)\n\n    if real_res == 0:\n        return 0\n    else:\n        return cutoff_res / real_res\n</code></pre>\n<pre><code>@njit()\ndef compute_weighted_gini(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    randoms = np.cumsum(weights / np.sum(weights))\n    total_pos = np.sum(y_true_label_sorted * weights)\n    cum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n    # compute final score\n    lorentz = cum_pos_found / total_pos\n    gini = (lorentz - randoms) * weights\n\n    return np.sum(gini)\n</code></pre>\n<pre><code>def custom_metric(y_true_label, y_pred_proba):\n    # compute normalized_weighted_gini score\n    tmp_0 = compute_weighted_gini(\n        y_true_label, y_pred_proba\n    )\n    tmp_1 = compute_weighted_gini(\n        y_true_label, y_true_label\n    )\n\n    if tmp_1 == 0:\n        g = 0\n    else:\n        g = tmp_0 / tmp_1\n\n    # compute top_four_percent_captured score\n    d = compute_top_four_percent_captured(\n        y_true_label, y_pred_proba\n    )\n\n    return 0.5 * (g + d)\n</code></pre>",
      "rawMarkdown": "Slight formatting issue with your discussion post, I believe you meant this:\n\n```\n@njit()\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    # compute cut-off point\n    weights_cumsum = np.cumsum(weights)\n    four_pct_cutoff = int(0.04 * np.sum(weights))\n\n    for i in range(len(weights)):\n        if weights_cumsum[i] > four_pct_cutoff:\n            break\n    y_true_label_sorted_cut = y_true_label_sorted[:i]\n\n    # compute final score\n    cutoff_res = np.sum(y_true_label_sorted_cut == 1)\n    real_res = np.sum(y_true_label_sorted == 1)\n\n    if real_res == 0:\n        return 0\n    else:\n        return cutoff_res / real_res\n```\n```\n@njit()\ndef compute_weighted_gini(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    randoms = np.cumsum(weights / np.sum(weights))\n    total_pos = np.sum(y_true_label_sorted * weights)\n    cum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n    # compute final score\n    lorentz = cum_pos_found / total_pos\n    gini = (lorentz - randoms) * weights\n\n    return np.sum(gini)\n```\n```\ndef custom_metric(y_true_label, y_pred_proba):\n    # compute normalized_weighted_gini score\n    tmp_0 = compute_weighted_gini(\n        y_true_label, y_pred_proba\n    )\n    tmp_1 = compute_weighted_gini(\n        y_true_label, y_true_label\n    )\n\n    if tmp_1 == 0:\n        g = 0\n    else:\n        g = tmp_0 / tmp_1\n\n    # compute top_four_percent_captured score\n    d = compute_top_four_percent_captured(\n        y_true_label, y_pred_proba\n    )\n\n    return 0.5 * (g + d)\n```",
      "votes": null
    },
    {
      "id": "1804572",
      "postDate": "05/29/2022 07:16:08",
      "content": "<p>Thanks ! That's extract what I mean.</p>",
      "rawMarkdown": "Thanks ! That's extract what I mean.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1803680,
      "author_name": "zhuoyin94",
      "author_url": "",
      "post_date": "05/28/2022 04:32:48",
      "content": "<p>A simple unit test case without assertion:</p>\n<p>def test_custom_metrics():</p>\n<pre><code># test 1\npred_df = pd.DataFrame(\n    {'prediction': np.random.random(2000)}\n)\ntrue_df = pd.DataFrame(\n    {'target': np.random.randint(0, 2, 2000)}\n)\n\nprint(amex_metric(true_df, pred_df))\n\nprint(\n    custom_metric(\n        true_df['target'].values, pred_df['prediction'].values\n    )\n)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1804318,
      "author_name": "munumbutt",
      "author_url": "",
      "post_date": "05/28/2022 18:59:10",
      "content": "<p>Slight formatting issue with your discussion post, I believe you meant this:</p>\n<pre><code>@njit()\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    # compute cut-off point\n    weights_cumsum = np.cumsum(weights)\n    four_pct_cutoff = int(0.04 * np.sum(weights))\n\n    for i in range(len(weights)):\n        if weights_cumsum[i] &gt; four_pct_cutoff:\n            break\n    y_true_label_sorted_cut = y_true_label_sorted[:i]\n\n    # compute final score\n    cutoff_res = np.sum(y_true_label_sorted_cut == 1)\n    real_res = np.sum(y_true_label_sorted == 1)\n\n    if real_res == 0:\n        return 0\n    else:\n        return cutoff_res / real_res\n</code></pre>\n<pre><code>@njit()\ndef compute_weighted_gini(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    randoms = np.cumsum(weights / np.sum(weights))\n    total_pos = np.sum(y_true_label_sorted * weights)\n    cum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n    # compute final score\n    lorentz = cum_pos_found / total_pos\n    gini = (lorentz - randoms) * weights\n\n    return np.sum(gini)\n</code></pre>\n<pre><code>def custom_metric(y_true_label, y_pred_proba):\n    # compute normalized_weighted_gini score\n    tmp_0 = compute_weighted_gini(\n        y_true_label, y_pred_proba\n    )\n    tmp_1 = compute_weighted_gini(\n        y_true_label, y_true_label\n    )\n\n    if tmp_1 == 0:\n        g = 0\n    else:\n        g = tmp_0 / tmp_1\n\n    # compute top_four_percent_captured score\n    d = compute_top_four_percent_captured(\n        y_true_label, y_pred_proba\n    )\n\n    return 0.5 * (g + d)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1804572,
          "author_name": "zhuoyin94",
          "author_url": "",
          "post_date": "05/29/2022 07:16:08",
          "content": "<p>Thanks ! That's extract what I mean.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1803678": "Since the original AMEX metric is achived using pandas, it is limited when you use is as an early stopping metric in XGBoost or Keras.\n\nI rewrite the metric using the numpy, and tested it with some unit test cases. Furthermore, the metric is accelerated by numba, which can make the computation much faster.\n\n@njit\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):\n\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    # compute cut-off point\n    weights_cumsum = np.cumsum(weights)\n    four_pct_cutoff = int(0.04 * np.sum(weights))\n\n    for i in range(len(weights)):\n        if weights_cumsum[i] > four_pct_cutoff:\n            break\n    y_true_label_sorted_cut = y_true_label_sorted[:i]\n\n    # compute final score\n    cutoff_res = np.sum(y_true_label_sorted_cut == 1)\n    real_res = np.sum(y_true_label_sorted == 1)\n\n    if real_res == 0:\n        return 0\n    else:\n        return cutoff_res / real_res\n\n\n@njit\ndef compute_weighted_gini(y_true_label, y_pred_proba):\n\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    randoms = np.cumsum(weights / np.sum(weights))\n    total_pos = np.sum(y_true_label_sorted * weights)\n    cum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n    # compute final score\n    lorentz = cum_pos_found / total_pos\n    gini = (lorentz - randoms) * weights\n\n    return np.sum(gini)\n\n\ndef custom_metric(y_true_label, y_pred_proba):\n\n    # compute normalized_weighted_gini score\n    tmp_0 = compute_weighted_gini(\n        y_true_label, y_pred_proba\n    )\n    tmp_1 = compute_weighted_gini(\n        y_true_label, y_true_label\n    )\n\n    if tmp_1 == 0:\n        g = 0\n    else:\n        g = tmp_0 / tmp_1\n\n    # compute top_four_percent_captured score\n    d = compute_top_four_percent_captured(\n        y_true_label, y_pred_proba\n    )\n\n    return 0.5 * (g + d)",
    "1803680": "A simple unit test case without assertion:\n\ndef test_custom_metrics():\n\n    # test 1\n    pred_df = pd.DataFrame(\n        {'prediction': np.random.random(2000)}\n    )\n    true_df = pd.DataFrame(\n        {'target': np.random.randint(0, 2, 2000)}\n    )\n\n    print(amex_metric(true_df, pred_df))\n\n    print(\n        custom_metric(\n            true_df['target'].values, pred_df['prediction'].values\n        )\n    )",
    "1804318": "Slight formatting issue with your discussion post, I believe you meant this:\n\n```\n@njit()\ndef compute_top_four_percent_captured(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    # compute cut-off point\n    weights_cumsum = np.cumsum(weights)\n    four_pct_cutoff = int(0.04 * np.sum(weights))\n\n    for i in range(len(weights)):\n        if weights_cumsum[i] > four_pct_cutoff:\n            break\n    y_true_label_sorted_cut = y_true_label_sorted[:i]\n\n    # compute final score\n    cutoff_res = np.sum(y_true_label_sorted_cut == 1)\n    real_res = np.sum(y_true_label_sorted == 1)\n\n    if real_res == 0:\n        return 0\n    else:\n        return cutoff_res / real_res\n```\n```\n@njit()\ndef compute_weighted_gini(y_true_label, y_pred_proba):\n    pred_sorted_idx = np.argsort(y_pred_proba)[::-1]\n    y_true_label_sorted = y_true_label[pred_sorted_idx]\n\n    weights = np.ones((len(y_true_label_sorted), ))\n    for i in range(len(weights)):\n        if y_true_label_sorted[i] == 0:\n            weights[i] = 20\n\n    randoms = np.cumsum(weights / np.sum(weights))\n    total_pos = np.sum(y_true_label_sorted * weights)\n    cum_pos_found = np.cumsum(y_true_label_sorted * weights)\n\n    # compute final score\n    lorentz = cum_pos_found / total_pos\n    gini = (lorentz - randoms) * weights\n\n    return np.sum(gini)\n```\n```\ndef custom_metric(y_true_label, y_pred_proba):\n    # compute normalized_weighted_gini score\n    tmp_0 = compute_weighted_gini(\n        y_true_label, y_pred_proba\n    )\n    tmp_1 = compute_weighted_gini(\n        y_true_label, y_true_label\n    )\n\n    if tmp_1 == 0:\n        g = 0\n    else:\n        g = tmp_0 / tmp_1\n\n    # compute top_four_percent_captured score\n    d = compute_top_four_percent_captured(\n        y_true_label, y_pred_proba\n    )\n\n    return 0.5 * (g + d)\n```",
    "1804572": "Thanks ! That's extract what I mean."
  },
  "source": "meta"
}