{
  "id": 82591,
  "title": "matthews correlation with a trageted balance",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/82591",
  "author_name": "",
  "post_date": "2019-03-02T14:46:42.656270900Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am working on robust cross validation techniques using the standard StratifiedKFold as in most kernels. This relies on about 6% positive samples in the training dataset. However, my models predict less than 4% positive samples over the test data set. This is kind of inconsistent especially if matthews correlation is used for early stopping and for calculating the CV score.</p>\n\n<p>Doesn't make any sense to keep the CV based on StratifiedKFold and only adjust the samples weights in the matthews correlation function as if we have a specified targeted positive samples percent. I edited the code below and would like to get feedback? BTW, after using this, my local cv scores (10 folds) are still not correlated to with the public LB !!</p>\n\n<p>```\ndef matthews_correlation_np(y_true, y_pred, threshold=0.5):\n    y_pred_pos = (y_pred &gt; threshold).astype(np.int64)\n    y_pred_neg = (1 - y_pred_pos)\n    y_pos = (y_true &gt; threshold).astype(np.int64)\n    y_neg = (1 - y_pos)</p>\n\n<pre><code>n_pos = np.sum(y_pos)\nn_neg = np.sum(y_neg)\n\n# desired_balance = n_pos / (n_pos + weight_neg * n_neg)\ndesired_balance = 0.04\nweight_neg = (n_pos / desired_balance - n_pos) / n_neg\ntp = np.sum(y_pos * y_pred_pos)\ntn = weight_neg * np.sum(y_neg * y_pred_neg)\nfp = weight_neg * np.sum(y_neg * y_pred_pos)\nfn = np.sum(y_pos * y_pred_neg)\n\nnumerator = (tp * tn - fp * fn)\ndenominator = np.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\nreturn numerator / (denominator + np.finfo(np.float64).eps)\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "482227",
      "postDate": "03/02/2019 14:46:42",
      "content": "<p>I am working on robust cross validation techniques using the standard StratifiedKFold as in most kernels. This relies on about 6% positive samples in the training dataset. However, my models predict less than 4% positive samples over the test data set. This is kind of inconsistent especially if matthews correlation is used for early stopping and for calculating the CV score.</p>\n\n<p>Doesn't make any sense to keep the CV based on StratifiedKFold and only adjust the samples weights in the matthews correlation function as if we have a specified targeted positive samples percent. I edited the code below and would like to get feedback? BTW, after using this, my local cv scores (10 folds) are still not correlated to with the public LB !!</p>\n\n<p>```\ndef matthews_correlation_np(y_true, y_pred, threshold=0.5):\n    y_pred_pos = (y_pred &gt; threshold).astype(np.int64)\n    y_pred_neg = (1 - y_pred_pos)\n    y_pos = (y_true &gt; threshold).astype(np.int64)\n    y_neg = (1 - y_pos)</p>\n\n<pre><code>n_pos = np.sum(y_pos)\nn_neg = np.sum(y_neg)\n\n# desired_balance = n_pos / (n_pos + weight_neg * n_neg)\ndesired_balance = 0.04\nweight_neg = (n_pos / desired_balance - n_pos) / n_neg\ntp = np.sum(y_pos * y_pred_pos)\ntn = weight_neg * np.sum(y_neg * y_pred_neg)\nfp = weight_neg * np.sum(y_neg * y_pred_pos)\nfn = np.sum(y_pos * y_pred_neg)\n\nnumerator = (tp * tn - fp * fn)\ndenominator = np.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\nreturn numerator / (denominator + np.finfo(np.float64).eps)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I am working on robust cross validation techniques using the standard StratifiedKFold as in most kernels. This relies on about 6% positive samples in the training dataset. However, my models predict less than 4% positive samples over the test data set. This is kind of inconsistent especially if matthews correlation is used for early stopping and for calculating the CV score.\n\nDoesn't make any sense to keep the CV based on StratifiedKFold and only adjust the samples weights in the matthews correlation function as if we have a specified targeted positive samples percent. I edited the code below and would like to get feedback? BTW, after using this, my local cv scores (10 folds) are still not correlated to with the public LB !!\n\n```\ndef matthews_correlation_np(y_true, y_pred, threshold=0.5):\n    y_pred_pos = (y_pred &gt; threshold).astype(np.int64)\n    y_pred_neg = (1 - y_pred_pos)\n    y_pos = (y_true &gt; threshold).astype(np.int64)\n    y_neg = (1 - y_pos)\n\n    n_pos = np.sum(y_pos)\n    n_neg = np.sum(y_neg)\n\n    # desired_balance = n_pos / (n_pos + weight_neg * n_neg)\n    desired_balance = 0.04\n    weight_neg = (n_pos / desired_balance - n_pos) / n_neg\n    tp = np.sum(y_pos * y_pred_pos)\n    tn = weight_neg * np.sum(y_neg * y_pred_neg)\n    fp = weight_neg * np.sum(y_neg * y_pred_pos)\n    fn = np.sum(y_pos * y_pred_neg)\n\n    numerator = (tp * tn - fp * fn)\n    denominator = np.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\n    return numerator / (denominator + np.finfo(np.float64).eps)\n\n```",
      "votes": null
    },
    {
      "id": "482565",
      "postDate": "03/03/2019 08:46:26",
      "content": "<p>I was just thinking about the same thing, same formula. Conversely, what is the positive fraction that matches your LB? I speculated that the positive sample is 2% in the test set, but that doesn't completely explain my LB scores. Probably there is something different other than the positive fraction...</p>",
      "rawMarkdown": "I was just thinking about the same thing, same formula. Conversely, what is the positive fraction that matches your LB? I speculated that the positive sample is 2% in the test set, but that doesn't completely explain my LB scores. Probably there is something different other than the positive fraction...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 482565,
      "author_name": "junkoda",
      "author_url": "",
      "post_date": "03/03/2019 08:46:26",
      "content": "<p>I was just thinking about the same thing, same formula. Conversely, what is the positive fraction that matches your LB? I speculated that the positive sample is 2% in the test set, but that doesn't completely explain my LB scores. Probably there is something different other than the positive fraction...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "482227": "I am working on robust cross validation techniques using the standard StratifiedKFold as in most kernels. This relies on about 6% positive samples in the training dataset. However, my models predict less than 4% positive samples over the test data set. This is kind of inconsistent especially if matthews correlation is used for early stopping and for calculating the CV score.\n\nDoesn't make any sense to keep the CV based on StratifiedKFold and only adjust the samples weights in the matthews correlation function as if we have a specified targeted positive samples percent. I edited the code below and would like to get feedback? BTW, after using this, my local cv scores (10 folds) are still not correlated to with the public LB !!\n\n```\ndef matthews_correlation_np(y_true, y_pred, threshold=0.5):\n    y_pred_pos = (y_pred &gt; threshold).astype(np.int64)\n    y_pred_neg = (1 - y_pred_pos)\n    y_pos = (y_true &gt; threshold).astype(np.int64)\n    y_neg = (1 - y_pos)\n\n    n_pos = np.sum(y_pos)\n    n_neg = np.sum(y_neg)\n\n    # desired_balance = n_pos / (n_pos + weight_neg * n_neg)\n    desired_balance = 0.04\n    weight_neg = (n_pos / desired_balance - n_pos) / n_neg\n    tp = np.sum(y_pos * y_pred_pos)\n    tn = weight_neg * np.sum(y_neg * y_pred_neg)\n    fp = weight_neg * np.sum(y_neg * y_pred_pos)\n    fn = np.sum(y_pos * y_pred_neg)\n\n    numerator = (tp * tn - fp * fn)\n    denominator = np.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\n    return numerator / (denominator + np.finfo(np.float64).eps)\n\n```",
    "482565": "I was just thinking about the same thing, same formula. Conversely, what is the positive fraction that matches your LB? I speculated that the positive sample is 2% in the test set, but that doesn't completely explain my LB scores. Probably there is something different other than the positive fraction..."
  },
  "source": "meta"
}