{
  "id": 341553,
  "title": "competition metric running on the GPU [pytorch]",
  "url": "/competitions/amex-default-prediction/discussion/341553",
  "author_name": "Radek Osmulski",
  "post_date": "2022-08-03T11:06:32.292000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I know I am a little bit late to the party, but just wanted to share with you the metric implemented in pytorch:</p>\n<pre><code># https://www.kaggle.com/kyakovlev\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327534\ndef amex_metric_mod_torch(y_true, y_pred):\n\n    y_true = torch.tensor(y_true).cuda()\n    y_pred = torch.tensor(y_pred).cuda()\n    labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n    labels = labels[labels[:, 1].argsort(descending=True)]\n\n    weights    = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n    cut_vals   = labels[torch.cumsum(weights, 0) &lt;= int(0.04 * torch.sum(weights))]\n    top_four   = torch.sum(cut_vals[:,0]) / torch.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n        labels = labels[labels[:, i].argsort(descending=True)]\n        weight         = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n        weight_random  = torch.cumsum(weight / torch.sum(weight), 0)\n        total_pos      = torch.sum(labels[:, 0] *  weight)\n        cum_pos_found  = torch.cumsum(labels[:, 0] * weight, 0)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = torch.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four).detach().cpu().item()\n</code></pre>\n<p>This should offer quite a speed up 🙂 Happy kaggling!</p>",
  "messages": [
    {
      "id": 1882622,
      "postDate": "2022-08-03T11:06:32.293Z",
      "content": "<p>Hey!</p>\n<p>I know I am a little bit late to the party, but just wanted to share with you the metric implemented in pytorch:</p>\n<pre><code># https://www.kaggle.com/kyakovlev\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327534\ndef amex_metric_mod_torch(y_true, y_pred):\n\n    y_true = torch.tensor(y_true).cuda()\n    y_pred = torch.tensor(y_pred).cuda()\n    labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n    labels = labels[labels[:, 1].argsort(descending=True)]\n\n    weights    = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n    cut_vals   = labels[torch.cumsum(weights, 0) &lt;= int(0.04 * torch.sum(weights))]\n    top_four   = torch.sum(cut_vals[:,0]) / torch.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n        labels = labels[labels[:, i].argsort(descending=True)]\n        weight         = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n        weight_random  = torch.cumsum(weight / torch.sum(weight), 0)\n        total_pos      = torch.sum(labels[:, 0] *  weight)\n        cum_pos_found  = torch.cumsum(labels[:, 0] * weight, 0)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = torch.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four).detach().cpu().item()\n</code></pre>\n<p>This should offer quite a speed up 🙂 Happy kaggling!</p>",
      "rawMarkdown": "Hey!\n\nI know I am a little bit late to the party, but just wanted to share with you the metric implemented in pytorch:\n\n```\n# https://www.kaggle.com/kyakovlev\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327534\ndef amex_metric_mod_torch(y_true, y_pred):\n    \n    y_true = torch.tensor(y_true).cuda()\n    y_pred = torch.tensor(y_pred).cuda()\n    labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n    labels = labels[labels[:, 1].argsort(descending=True)]\n\n    weights    = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n    cut_vals   = labels[torch.cumsum(weights, 0) <= int(0.04 * torch.sum(weights))]\n    top_four   = torch.sum(cut_vals[:,0]) / torch.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n        labels = labels[labels[:, i].argsort(descending=True)]\n        weight         = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n        weight_random  = torch.cumsum(weight / torch.sum(weight), 0)\n        total_pos      = torch.sum(labels[:, 0] *  weight)\n        cum_pos_found  = torch.cumsum(labels[:, 0] * weight, 0)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = torch.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four).detach().cpu().item()\n```\n\nThis should offer quite a speed up 🙂 Happy kaggling!",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1882622": "Hey!\n\nI know I am a little bit late to the party, but just wanted to share with you the metric implemented in pytorch:\n\n```\n# https://www.kaggle.com/kyakovlev\n# https://www.kaggle.com/competitions/amex-default-prediction/discussion/327534\ndef amex_metric_mod_torch(y_true, y_pred):\n    \n    y_true = torch.tensor(y_true).cuda()\n    y_pred = torch.tensor(y_pred).cuda()\n    labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n    labels = labels[labels[:, 1].argsort(descending=True)]\n\n    weights    = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n    cut_vals   = labels[torch.cumsum(weights, 0) <= int(0.04 * torch.sum(weights))]\n    top_four   = torch.sum(cut_vals[:,0]) / torch.sum(labels[:,0])\n\n    gini = [0,0]\n    for i in [1,0]:\n        labels = torch.cat([y_true.reshape(-1,1), y_pred.reshape(-1,1)], axis=1)\n        labels = labels[labels[:, i].argsort(descending=True)]\n        weight         = torch.tensor(np.where(labels[:,0].cpu()==0, 20, 1)).cuda()\n        weight_random  = torch.cumsum(weight / torch.sum(weight), 0)\n        total_pos      = torch.sum(labels[:, 0] *  weight)\n        cum_pos_found  = torch.cumsum(labels[:, 0] * weight, 0)\n        lorentz        = cum_pos_found / total_pos\n        gini[i]        = torch.sum((lorentz - weight_random) * weight)\n\n    return 0.5 * (gini[1]/gini[0] + top_four).detach().cpu().item()\n```\n\nThis should offer quite a speed up 🙂 Happy kaggling!"
  }
}