{
  "id": 79038,
  "title": "Fast Matthews Correlation computation (CPU)",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/79038",
  "author_name": "",
  "post_date": "2019-01-30T08:14:41.319373Z",
  "votes": 23,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was doing a few tests and could not believe TF Mathiews Correlation took so long when computing scores for different thresholds, so I decided to code my own version (I did not check all  kernels, so maybe it's already there somewhere).</p>\n\n<p>Here we go (numpy is quick):</p>\n\n<pre><code>def my_mat_cor(y_true, y_pred):\n    assert y_true.shape[0] == y_pred.shape[0]\n\n    tp = np.sum((y_true == 1) &amp;amp; (y_pred == 1))\n    tn = np.sum((y_true == 0) &amp;amp; (y_pred == 0))\n    fp = np.sum((y_true == 0) &amp;amp; (y_pred == 1))\n    fn = np.sum((y_true == 1) &amp;amp; (y_pred == 0))\n\n    numerator = (tp * tn - fp * fn) \n    denominator = ((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn)) ** .5\n\n    return numerator / (denominator + 1e-15)\n</code></pre>\n\n<p>This call takes 4ms in a kernel, while TF mathews correlation function used in most of the kernels take 2s.</p>\n\n<p>I believe you want to use the above function when you compute scores for different thresholds.</p>\n\n<p>When you compute the best threshold I guess you want to use CPMP's function <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682\">here</a>.</p>\n\n<p>Hope this helps :)</p>\n\n<p>UPDATE:\nI saw sklearn offered a matthew correlation metric :) So here is a benchmark</p>",
  "messages": [
    {
      "id": "463576",
      "postDate": "01/30/2019 08:14:41",
      "content": "<p>I was doing a few tests and could not believe TF Mathiews Correlation took so long when computing scores for different thresholds, so I decided to code my own version (I did not check all  kernels, so maybe it's already there somewhere).</p>\n\n<p>Here we go (numpy is quick):</p>\n\n<pre><code>def my_mat_cor(y_true, y_pred):\n    assert y_true.shape[0] == y_pred.shape[0]\n\n    tp = np.sum((y_true == 1) &amp;amp; (y_pred == 1))\n    tn = np.sum((y_true == 0) &amp;amp; (y_pred == 0))\n    fp = np.sum((y_true == 0) &amp;amp; (y_pred == 1))\n    fn = np.sum((y_true == 1) &amp;amp; (y_pred == 0))\n\n    numerator = (tp * tn - fp * fn) \n    denominator = ((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn)) ** .5\n\n    return numerator / (denominator + 1e-15)\n</code></pre>\n\n<p>This call takes 4ms in a kernel, while TF mathews correlation function used in most of the kernels take 2s.</p>\n\n<p>I believe you want to use the above function when you compute scores for different thresholds.</p>\n\n<p>When you compute the best threshold I guess you want to use CPMP's function <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682\">here</a>.</p>\n\n<p>Hope this helps :)</p>\n\n<p>UPDATE:\nI saw sklearn offered a matthew correlation metric :) So here is a benchmark</p>",
      "rawMarkdown": "I was doing a few tests and could not believe TF Mathiews Correlation took so long when computing scores for different thresholds, so I decided to code my own version (I did not check all  kernels, so maybe it's already there somewhere).\n\nHere we go (numpy is quick):\n\n    def my_mat_cor(y_true, y_pred):\n        assert y_true.shape[0] == y_pred.shape[0]\n    \n        tp = np.sum((y_true == 1) &amp; (y_pred == 1))\n        tn = np.sum((y_true == 0) &amp; (y_pred == 0))\n        fp = np.sum((y_true == 0) &amp; (y_pred == 1))\n        fn = np.sum((y_true == 1) &amp; (y_pred == 0))\n    \n        numerator = (tp * tn - fp * fn) \n        denominator = ((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn)) ** .5\n\n        return numerator / (denominator + 1e-15)\n\nThis call takes 4ms in a kernel, while TF mathews correlation function used in most of the kernels take 2s.\n\nI believe you want to use the above function when you compute scores for different thresholds.\n\nWhen you compute the best threshold I guess you want to use CPMP's function [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682).\n\nHope this helps :)\n\nUPDATE:\nI saw sklearn offered a matthew correlation metric :) So here is a benchmark",
      "votes": null
    },
    {
      "id": "464080",
      "postDate": "01/31/2019 06:55:48",
      "content": "<p>Thank you and welcome to VSB Power Line Fault Detection.\nAs the leaderboard spread shows, this one is better than Microsoft Malware Prediction  :)</p>",
      "rawMarkdown": "Thank you and welcome to VSB Power Line Fault Detection.\nAs the leaderboard spread shows, this one is better than Microsoft Malware Prediction  :)",
      "votes": null
    },
    {
      "id": "464085",
      "postDate": "01/31/2019 07:10:44",
      "content": "<p>Thanks ;-) Not sure about that really. I have yet to see a correlation between CV and LB !</p>\n\n<p>I've only done a few tests and this one very looks like Mercedes competition for me right now. </p>\n\n<p>Competitions with threshold optimization are particularly difficult IMHO.</p>",
      "rawMarkdown": "Thanks ;-) Not sure about that really. I have yet to see a correlation between CV and LB !\n\nI've only done a few tests and this one very looks like Mercedes competition for me right now. \n\nCompetitions with threshold optimization are particularly difficult IMHO.",
      "votes": null
    },
    {
      "id": "464601",
      "postDate": "02/01/2019 06:27:58",
      "content": "<p>agree. sklearn's function is super slow </p>\n\n<p><a href=\"https://github.com/kmedian/korr/blob/master/examples/mcc%20(Matthews%20correlation).ipynb\">https://github.com/kmedian/korr/blob/master/examples/mcc%20(Matthews%20correlation).ipynb</a></p>\n\n<p>for tiny/small datasets it does not matter. for a millions of examples plain numpy is somewhat 16x times faster. sklearn must have something that screw their big-O.</p>\n\n<p><a href=\"https://github.com/kmedian/korr/blob/master/korr/confusion.py\">korr.confusion</a>\n<a href=\"https://github.com/kmedian/korr/blob/master/korr/mcc.py\">korr.mcc</a></p>",
      "rawMarkdown": "agree. sklearn's function is super slow \n\nhttps://github.com/kmedian/korr/blob/master/examples/mcc%20(Matthews%20correlation).ipynb\n\nfor tiny/small datasets it does not matter. for a millions of examples plain numpy is somewhat 16x times faster. sklearn must have something that screw their big-O.\n\n[korr.confusion](https://github.com/kmedian/korr/blob/master/korr/confusion.py)\n[korr.mcc](https://github.com/kmedian/korr/blob/master/korr/mcc.py)",
      "votes": null
    },
    {
      "id": "465517",
      "postDate": "02/03/2019 10:35:09",
      "content": "<p>Yes, experiencing that a lot now in the last few days. While the CV is increasing the LB score is decreasing. Unable to see how to proceed further.  Also there are lot of fluctuations in mcc during training.</p>\n\n<p>Can we not stick with a hard coded threshold like 0.5 and try to train a model. That was my initial approach, but after seeing the thresold optimization code I added that to my model.</p>",
      "rawMarkdown": "Yes, experiencing that a lot now in the last few days. While the CV is increasing the LB score is decreasing. Unable to see how to proceed further.  Also there are lot of fluctuations in mcc during training.\n\nCan we not stick with a hard coded threshold like 0.5 and try to train a model. That was my initial approach, but after seeing the thresold optimization code I added that to my model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 464080,
      "author_name": "subrahmanyamv",
      "author_url": "",
      "post_date": "01/31/2019 06:55:48",
      "content": "<p>Thank you and welcome to VSB Power Line Fault Detection.\nAs the leaderboard spread shows, this one is better than Microsoft Malware Prediction  :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 464085,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "01/31/2019 07:10:44",
          "content": "<p>Thanks ;-) Not sure about that really. I have yet to see a correlation between CV and LB !</p>\n\n<p>I've only done a few tests and this one very looks like Mercedes competition for me right now. </p>\n\n<p>Competitions with threshold optimization are particularly difficult IMHO.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 465517,
          "author_name": "subrahmanyamv",
          "author_url": "",
          "post_date": "02/03/2019 10:35:09",
          "content": "<p>Yes, experiencing that a lot now in the last few days. While the CV is increasing the LB score is decreasing. Unable to see how to proceed further.  Also there are lot of fluctuations in mcc during training.</p>\n\n<p>Can we not stick with a hard coded threshold like 0.5 and try to train a model. That was my initial approach, but after seeing the thresold optimization code I added that to my model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 464601,
      "author_name": "bubblride",
      "author_url": "",
      "post_date": "02/01/2019 06:27:58",
      "content": "<p>agree. sklearn's function is super slow </p>\n\n<p><a href=\"https://github.com/kmedian/korr/blob/master/examples/mcc%20(Matthews%20correlation).ipynb\">https://github.com/kmedian/korr/blob/master/examples/mcc%20(Matthews%20correlation).ipynb</a></p>\n\n<p>for tiny/small datasets it does not matter. for a millions of examples plain numpy is somewhat 16x times faster. sklearn must have something that screw their big-O.</p>\n\n<p><a href=\"https://github.com/kmedian/korr/blob/master/korr/confusion.py\">korr.confusion</a>\n<a href=\"https://github.com/kmedian/korr/blob/master/korr/mcc.py\">korr.mcc</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "463576": "I was doing a few tests and could not believe TF Mathiews Correlation took so long when computing scores for different thresholds, so I decided to code my own version (I did not check all  kernels, so maybe it's already there somewhere).\n\nHere we go (numpy is quick):\n\n    def my_mat_cor(y_true, y_pred):\n        assert y_true.shape[0] == y_pred.shape[0]\n    \n        tp = np.sum((y_true == 1) &amp; (y_pred == 1))\n        tn = np.sum((y_true == 0) &amp; (y_pred == 0))\n        fp = np.sum((y_true == 0) &amp; (y_pred == 1))\n        fn = np.sum((y_true == 1) &amp; (y_pred == 0))\n    \n        numerator = (tp * tn - fp * fn) \n        denominator = ((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn)) ** .5\n\n        return numerator / (denominator + 1e-15)\n\nThis call takes 4ms in a kernel, while TF mathews correlation function used in most of the kernels take 2s.\n\nI believe you want to use the above function when you compute scores for different thresholds.\n\nWhen you compute the best threshold I guess you want to use CPMP's function [here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682).\n\nHope this helps :)\n\nUPDATE:\nI saw sklearn offered a matthew correlation metric :) So here is a benchmark",
    "464080": "Thank you and welcome to VSB Power Line Fault Detection.\nAs the leaderboard spread shows, this one is better than Microsoft Malware Prediction  :)",
    "464085": "Thanks ;-) Not sure about that really. I have yet to see a correlation between CV and LB !\n\nI've only done a few tests and this one very looks like Mercedes competition for me right now. \n\nCompetitions with threshold optimization are particularly difficult IMHO.",
    "464601": "agree. sklearn's function is super slow \n\nhttps://github.com/kmedian/korr/blob/master/examples/mcc%20(Matthews%20correlation).ipynb\n\nfor tiny/small datasets it does not matter. for a millions of examples plain numpy is somewhat 16x times faster. sklearn must have something that screw their big-O.\n\n[korr.confusion](https://github.com/kmedian/korr/blob/master/korr/confusion.py)\n[korr.mcc](https://github.com/kmedian/korr/blob/master/korr/mcc.py)",
    "465517": "Yes, experiencing that a lot now in the last few days. While the CV is increasing the LB score is decreasing. Unable to see how to proceed further.  Also there are lot of fluctuations in mcc during training.\n\nCan we not stick with a hard coded threshold like 0.5 and try to train a model. That was my initial approach, but after seeing the thresold optimization code I added that to my model."
  },
  "source": "meta"
}