{
  "id": 327534,
  "title": "Metric without DF",
  "url": "/competitions/amex-default-prediction/discussion/327534",
  "author_name": "",
  "post_date": "2022-05-27T15:55:08.552252200Z",
  "votes": 73,
  "comment_count": 4,
  "views": 0,
  "content": "<p>UPDATE:<br>\nPlease use faster and cleaner metric function:<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020</a><br>\nby <a href=\"https://www.kaggle.com/yunchonggan\" target=\"_blank\">@yunchonggan</a> </p>\n<p>and here metric speed comparison <br>\n<a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\">https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations</a><br>\nby <a href=\"https://www.kaggle.com/rohanraogopal\" target=\"_blank\">@rohanraogopal</a> </p>\n<hr>\n<p>Old post</p>\n<p>Here is a competition metric without pandas and DataFrames.</p>\n<p>Ugly code works slightly faster (hope someone will make a normal one).</p>\n<p>It's a temporary topic and will be removed in a few days.</p>\n<p>Differences against all Zeros -&gt; come from sorting differences (argsort doesn't take into account initial index and can give bit different results if predictions contain \"shared\" values) -&gt; but it is 100% same on real predictions</p>\n<pre><code>def amex_metric_mod(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) &lt;= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n</code></pre>\n<pre><code>12.5 s ± 42.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n41.7 s ± 242 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\nnew -&gt; 0.2519002854103393\nold -&gt; 0.2519002854103393\n</code></pre>\n<p>Custom metric for LGBM</p>\n<pre><code>def amex_metric_mod_lgbm(y_pred: np.ndarray, data: lgb.Dataset):\n\n    y_true = data.get_label()\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) &lt;= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 'AMEX', 0.5 * (gini[1]/gini[0]+ top_four), True\n</code></pre>\n<p>usage:</p>\n<pre><code>estimator = lgb.train(lgb_params, train_data, \n                                    valid_sets = [train_data, valid_data],\n                                    verbose_eval = 100,\n                                    feval=amex_metric_mod_lgbm,\n                             )\n</code></pre>",
  "messages": [
    {
      "id": "1803248",
      "postDate": "05/27/2022 15:55:08",
      "content": "<p>UPDATE:<br>\nPlease use faster and cleaner metric function:<br>\n<a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328020</a><br>\nby <a href=\"https://www.kaggle.com/yunchonggan\" target=\"_blank\">@yunchonggan</a> </p>\n<p>and here metric speed comparison <br>\n<a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\">https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations</a><br>\nby <a href=\"https://www.kaggle.com/rohanraogopal\" target=\"_blank\">@rohanraogopal</a> </p>\n<hr>\n<p>Old post</p>\n<p>Here is a competition metric without pandas and DataFrames.</p>\n<p>Ugly code works slightly faster (hope someone will make a normal one).</p>\n<p>It's a temporary topic and will be removed in a few days.</p>\n<p>Differences against all Zeros -&gt; come from sorting differences (argsort doesn't take into account initial index and can give bit different results if predictions contain \"shared\" values) -&gt; but it is 100% same on real predictions</p>\n<pre><code>def amex_metric_mod(y_true, y_pred):\n\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) &lt;= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n</code></pre>\n<pre><code>12.5 s ± 42.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n41.7 s ± 242 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\nnew -&gt; 0.2519002854103393\nold -&gt; 0.2519002854103393\n</code></pre>\n<p>Custom metric for LGBM</p>\n<pre><code>def amex_metric_mod_lgbm(y_pred: np.ndarray, data: lgb.Dataset):\n\n    y_true = data.get_label()\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) &lt;= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 'AMEX', 0.5 * (gini[1]/gini[0]+ top_four), True\n</code></pre>\n<p>usage:</p>\n<pre><code>estimator = lgb.train(lgb_params, train_data, \n                                    valid_sets = [train_data, valid_data],\n                                    verbose_eval = 100,\n                                    feval=amex_metric_mod_lgbm,\n                             )\n</code></pre>",
      "rawMarkdown": "UPDATE:\nPlease use faster and cleaner metric function:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328020\nby @yunchonggan \n\nand here metric speed comparison \nhttps://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\nby @rohanraogopal \n\n\n----\n\nOld post\n\nHere is a competition metric without pandas and DataFrames.\n\nUgly code works slightly faster (hope someone will make a normal one).\n\nIt's a temporary topic and will be removed in a few days.\n\nDifferences against all Zeros -> come from sorting differences (argsort doesn't take into account initial index and can give bit different results if predictions contain \"shared\" values) -> but it is 100% same on real predictions\n\n```\ndef amex_metric_mod(y_true, y_pred):\n   \n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n```\n\n```\n12.5 s ± 42.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n41.7 s ± 242 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\nnew -> 0.2519002854103393\nold -> 0.2519002854103393\n```\n\n\nCustom metric for LGBM\n```\ndef amex_metric_mod_lgbm(y_pred: np.ndarray, data: lgb.Dataset):\n   \n    y_true = data.get_label()\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 'AMEX', 0.5 * (gini[1]/gini[0]+ top_four), True\n```\n\nusage:\n```\nestimator = lgb.train(lgb_params, train_data, \n                                    valid_sets = [train_data, valid_data],\n                                    verbose_eval = 100,\n                                    feval=amex_metric_mod_lgbm,\n                             )\n```",
      "votes": null
    },
    {
      "id": "1806639",
      "postDate": "05/31/2022 11:00:54",
      "content": "<p>Great, thanks!</p>",
      "rawMarkdown": "Great, thanks!",
      "votes": null
    },
    {
      "id": "1807920",
      "postDate": "06/01/2022 12:45:27",
      "content": "<p>I see that you are using only numpy to perform the process.<br>\nI'm learning a lot! Thank you! 🙌</p>",
      "rawMarkdown": "I see that you are using only numpy to perform the process.\nI'm learning a lot! Thank you! 🙌",
      "votes": null
    },
    {
      "id": "2680160",
      "postDate": "03/04/2024 02:34:56",
      "content": "<p>good work! thanks</p>",
      "rawMarkdown": "good work! thanks",
      "votes": null
    },
    {
      "id": "3162118",
      "postDate": "03/28/2025 20:10:00",
      "content": "<p>Great, thansk1!</p>",
      "rawMarkdown": "Great, thansk1!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1806639,
      "author_name": "chrgri",
      "author_url": "",
      "post_date": "05/31/2022 11:00:54",
      "content": "<p>Great, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1807920,
      "author_name": "tgwstr",
      "author_url": "",
      "post_date": "06/01/2022 12:45:27",
      "content": "<p>I see that you are using only numpy to perform the process.<br>\nI'm learning a lot! Thank you! 🙌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2680160,
      "author_name": "ddoooongdfdf",
      "author_url": "",
      "post_date": "03/04/2024 02:34:56",
      "content": "<p>good work! thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3162118,
      "author_name": "benzonsalazar",
      "author_url": "",
      "post_date": "03/28/2025 20:10:00",
      "content": "<p>Great, thansk1!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1803248": "UPDATE:\nPlease use faster and cleaner metric function:\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328020\nby @yunchonggan \n\nand here metric speed comparison \nhttps://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\nby @rohanraogopal \n\n\n----\n\nOld post\n\nHere is a competition metric without pandas and DataFrames.\n\nUgly code works slightly faster (hope someone will make a normal one).\n\nIt's a temporary topic and will be removed in a few days.\n\nDifferences against all Zeros -> come from sorting differences (argsort doesn't take into account initial index and can give bit different results if predictions contain \"shared\" values) -> but it is 100% same on real predictions\n\n```\ndef amex_metric_mod(y_true, y_pred):\n   \n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four)\n```\n\n```\n12.5 s ± 42.6 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\n41.7 s ± 242 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)\nnew -> 0.2519002854103393\nold -> 0.2519002854103393\n```\n\n\nCustom metric for LGBM\n```\ndef amex_metric_mod_lgbm(y_pred: np.ndarray, data: lgb.Dataset):\n   \n    y_true = data.get_label()\n    labels     = np.transpose(np.array([y_true, y_pred]))\n    labels     = labels[labels[:, 1].argsort()[::-1]]\n    weights    = np.where(labels[:,0]==0, 20, 1)\n    cut_vals   = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four   = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels         = np.transpose(np.array([y_true, y_pred]))\n        labels         = labels[labels[:, i].argsort()[::-1]]\n        weight         = np.where(labels[:,0]==0, 20, 1)\n        weight_random  = np.cumsum(weight / np.sum(weight))\n        total_pos      = np.sum(labels[:, 0] *  weight)\n        cum_pos_found  = np.cumsum(labels[:, 0] * weight)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = np.sum((lorentz - weight_random) * weight)\n\n    return 'AMEX', 0.5 * (gini[1]/gini[0]+ top_four), True\n```\n\nusage:\n```\nestimator = lgb.train(lgb_params, train_data, \n                                    valid_sets = [train_data, valid_data],\n                                    verbose_eval = 100,\n                                    feval=amex_metric_mod_lgbm,\n                             )\n```",
    "1806639": "Great, thanks!",
    "1807920": "I see that you are using only numpy to perform the process.\nI'm learning a lot! Thank you! 🙌",
    "2680160": "good work! thanks",
    "3162118": "Great, thansk1!"
  },
  "source": "meta"
}