{
  "id": 85212,
  "title": "Final standings secret",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85212",
  "author_name": "Andrii Sydorchuk",
  "post_date": "2019-03-22T09:37:54.003000",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>```\nimport numpy as np</p>\n\n<p>def final_standings(participants, solutions):\n    np.random.shuffle(participants)\n    return participants\n```</p>\n\n<p>Jokes aside, I am a bit disappointed with quality of data in this competition. It seems like final standings are mainly random.</p>\n\n<p>Nevertheless congrats to all winners!</p>",
  "messages": [
    {
      "id": 496510,
      "postDate": "2019-03-22T09:37:54.003Z",
      "content": "<p>```\nimport numpy as np</p>\n\n<p>def final_standings(participants, solutions):\n    np.random.shuffle(participants)\n    return participants\n```</p>\n\n<p>Jokes aside, I am a bit disappointed with quality of data in this competition. It seems like final standings are mainly random.</p>\n\n<p>Nevertheless congrats to all winners!</p>",
      "rawMarkdown": "```\nimport numpy as np\n\ndef final_standings(participants, solutions):\n    np.random.shuffle(participants)\n    return participants\n```\n\nJokes aside, I am a bit disappointed with quality of data in this competition. It seems like final standings are mainly random.\n\nNevertheless congrats to all winners!",
      "votes": 3
    },
    {
      "id": 496670,
      "postDate": "2019-03-22T13:03:23.730Z",
      "content": "<p>A lot of it is down to the metric rather than the data. Other Kaggle competitions that used MCC to evaluate submissions had big shake-ups too.</p>",
      "rawMarkdown": "A lot of it is down to the metric rather than the data. Other Kaggle competitions that used MCC to evaluate submissions had big shake-ups too.",
      "votes": 4,
      "replies": [
        {
          "id": 496764,
          "postDate": "2019-03-22T15:00:06.463Z",
          "content": "<p>It seems to be a mix of a small dataset with many features and a bad metric. What are the other competitions that used MCC?</p>",
          "rawMarkdown": "It seems to be a mix of a small dataset with many features and a bad metric. What are the other competitions that used MCC?",
          "votes": 1
        },
        {
          "id": 496770,
          "postDate": "2019-03-22T15:08:01.043Z",
          "content": "<p>IMHO, I think MCC itself is not a bad metric since it handles the imbalance class distribution quite well, \n<a href=\"https://en.wikipedia.org/wiki/Matthews_correlation_coefficient#Advantages_of_MCC_over_accuracy_and_F1_score\">https://en.wikipedia.org/wiki/Matthews_correlation_coefficient#Advantages_of_MCC_over_accuracy_and_F1_score</a></p>\n\n<p>Having said that, MCC score can be quite sensitive : only few more correct/wrong prediction can increase/decrease score a lot (if you are interested in this point, please take a look on my kernel on MCC heuristic)</p>",
          "rawMarkdown": "IMHO, I think MCC itself is not a bad metric since it handles the imbalance class distribution quite well, \nhttps://en.wikipedia.org/wiki/Matthews_correlation_coefficient#Advantages_of_MCC_over_accuracy_and_F1_score\n\nHaving said that, MCC score can be quite sensitive : only few more correct/wrong prediction can increase/decrease score a lot (if you are interested in this point, please take a look on my kernel on MCC heuristic)",
          "votes": 2
        },
        {
          "id": 496782,
          "postDate": "2019-03-22T15:17:40.830Z",
          "content": "<p>Indeed, it would be better to use a continuous version of MCC. I used it to experiment with models, as the results are more stable and easier to interpret.</p>",
          "rawMarkdown": "Indeed, it would be better to use a continuous version of MCC. I used it to experiment with models, as the results are more stable and easier to interpret.",
          "votes": 1
        },
        {
          "id": 496798,
          "postDate": "2019-03-22T15:36:54.203Z",
          "content": "<p>That's Interesting!  Could you please share a formula of its continuous version ?</p>",
          "rawMarkdown": "That's Interesting!  Could you please share a formula of its continuous version ?"
        },
        {
          "id": 496875,
          "postDate": "2019-03-22T17:37:48.633Z",
          "content": "<p>Do you think ROC AUC would have been a better metric?</p>",
          "rawMarkdown": "Do you think ROC AUC would have been a better metric?"
        },
        {
          "id": 496888,
          "postDate": "2019-03-22T17:56:26.960Z",
          "content": "<p>Here is continuous version (same as in kernels, but no rounding for <code>y_pred</code> variable):\n```\ndef matthews_correlation(y_true, y_pred):\n    y_pred_pos = K.clip(y_pred, 0, 1) # no rounding here!!!\n    y_pred_neg = 1 - y_pred_pos</p>\n\n<pre><code>y_pos = K.round(K.clip(y_true, 0, 1))\ny_neg = 1 - y_pos\n\ntp = K.sum(y_pos * y_pred_pos)\ntn = K.sum(y_neg * y_pred_neg)\n\nfp = K.sum(y_neg * y_pred_pos)\nfn = K.sum(y_pos * y_pred_neg)\n\nnumerator = (tp * tn - fp * fn)\ndenominator = K.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\n\nreturn numerator / (denominator + K.epsilon())\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "Here is continuous version (same as in kernels, but no rounding for `y_pred` variable):\n```\ndef matthews_correlation(y_true, y_pred):\n    y_pred_pos = K.clip(y_pred, 0, 1) # no rounding here!!!\n    y_pred_neg = 1 - y_pred_pos\n\n    y_pos = K.round(K.clip(y_true, 0, 1))\n    y_neg = 1 - y_pos\n\n    tp = K.sum(y_pos * y_pred_pos)\n    tn = K.sum(y_neg * y_pred_neg)\n\n    fp = K.sum(y_neg * y_pred_pos)\n    fn = K.sum(y_pos * y_pred_neg)\n\n    numerator = (tp * tn - fp * fn)\n    denominator = K.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\n\n    return numerator / (denominator + K.epsilon())\n```",
          "votes": 2
        },
        {
          "id": 496892,
          "postDate": "2019-03-22T17:58:54.890Z",
          "content": "<p>Jack, I am not sure ROC AUC would have been better, since it would require you to pick a threshold and a few samples on the wrong side of the threshold would change result significantly. So it has essentially the same downside as MCC when you have so few positive samples.</p>",
          "rawMarkdown": "Jack, I am not sure ROC AUC would have been better, since it would require you to pick a threshold and a few samples on the wrong side of the threshold would change result significantly. So it has essentially the same downside as MCC when you have so few positive samples.",
          "votes": 1
        },
        {
          "id": 496895,
          "postDate": "2019-03-22T18:06:52.370Z",
          "content": "<p>I think maybe clipping strategy may play a big role on the sensitivity of prediction results: <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/48241\">https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/48241</a></p>",
          "rawMarkdown": "I think maybe clipping strategy may play a big role on the sensitivity of prediction results: https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/48241",
          "votes": 13
        },
        {
          "id": 496936,
          "postDate": "2019-03-22T18:52:01.337Z",
          "content": "<p>Correct, the point was that competition results would be more stable if continuous metric was used for evaluation (e.g. RMSE, logloss, etc).</p>",
          "rawMarkdown": "Correct, the point was that competition results would be more stable if continuous metric was used for evaluation (e.g. RMSE, logloss, etc)."
        },
        {
          "id": 496988,
          "postDate": "2019-03-22T19:59:31.793Z",
          "content": "<p><a href=\"/asydorchuk\">@asydorchuk</a> what if it were like ROC AUC in the Santander competition were you  submit probabilities rather than 0 or 1? </p>",
          "rawMarkdown": "@asydorchuk what if it were like ROC AUC in the Santander competition were you  submit probabilities rather than 0 or 1? ",
          "votes": 1
        },
        {
          "id": 497017,
          "postDate": "2019-03-22T21:06:31.460Z",
          "content": "<p>Yes, that should improve evaluation stability.</p>",
          "rawMarkdown": "Yes, that should improve evaluation stability."
        },
        {
          "id": 497165,
          "postDate": "2019-03-23T03:54:21.627Z",
          "content": "<p>I had the same thought about submitting probabilities and scoring them by ROC AUC or some other continuous metric.  Although I'm not an expert on detecting power line faults, it seems to me that rounding predicted probabilities to 0 or 1 discards a lot of potentially useful information.  I wonder what the reasoning was behind using the MCC.</p>",
          "rawMarkdown": "I had the same thought about submitting probabilities and scoring them by ROC AUC or some other continuous metric.  Although I'm not an expert on detecting power line faults, it seems to me that rounding predicted probabilities to 0 or 1 discards a lot of potentially useful information.  I wonder what the reasoning was behind using the MCC.",
          "votes": 2
        }
      ]
    },
    {
      "id": 497379,
      "postDate": "2019-03-23T13:21:26.857Z",
      "content": "<p>Cool function! I will use it to predict my score in all future competitions :-)</p>",
      "rawMarkdown": "Cool function! I will use it to predict my score in all future competitions :-)\n",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 496670,
      "author_name": "Max Halford",
      "author_url": "",
      "post_date": "2019-03-22T13:03:23.730000",
      "content": "<p>A lot of it is down to the metric rather than the data. Other Kaggle competitions that used MCC to evaluate submissions had big shake-ups too.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 496764,
          "author_name": "Andrii Sydorchuk",
          "author_url": "",
          "post_date": "2019-03-22T15:00:06.463000",
          "content": "<p>It seems to be a mix of a small dataset with many features and a bad metric. What are the other competitions that used MCC?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496770,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-03-22T15:08:01.043000",
          "content": "<p>IMHO, I think MCC itself is not a bad metric since it handles the imbalance class distribution quite well, \n<a href=\"https://en.wikipedia.org/wiki/Matthews_correlation_coefficient#Advantages_of_MCC_over_accuracy_and_F1_score\">https://en.wikipedia.org/wiki/Matthews_correlation_coefficient#Advantages_of_MCC_over_accuracy_and_F1_score</a></p>\n\n<p>Having said that, MCC score can be quite sensitive : only few more correct/wrong prediction can increase/decrease score a lot (if you are interested in this point, please take a look on my kernel on MCC heuristic)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 496782,
          "author_name": "Andrii Sydorchuk",
          "author_url": "",
          "post_date": "2019-03-22T15:17:40.830000",
          "content": "<p>Indeed, it would be better to use a continuous version of MCC. I used it to experiment with models, as the results are more stable and easier to interpret.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496798,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-03-22T15:36:54.203000",
          "content": "<p>That's Interesting!  Could you please share a formula of its continuous version ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 496875,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-03-22T17:37:48.633000",
          "content": "<p>Do you think ROC AUC would have been a better metric?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 496888,
          "author_name": "Andrii Sydorchuk",
          "author_url": "",
          "post_date": "2019-03-22T17:56:26.960000",
          "content": "<p>Here is continuous version (same as in kernels, but no rounding for <code>y_pred</code> variable):\n```\ndef matthews_correlation(y_true, y_pred):\n    y_pred_pos = K.clip(y_pred, 0, 1) # no rounding here!!!\n    y_pred_neg = 1 - y_pred_pos</p>\n\n<pre><code>y_pos = K.round(K.clip(y_true, 0, 1))\ny_neg = 1 - y_pos\n\ntp = K.sum(y_pos * y_pred_pos)\ntn = K.sum(y_neg * y_pred_neg)\n\nfp = K.sum(y_neg * y_pred_pos)\nfn = K.sum(y_pos * y_pred_neg)\n\nnumerator = (tp * tn - fp * fn)\ndenominator = K.sqrt((tp + fp) * (tp + fn) * (tn + fp) * (tn + fn))\n\nreturn numerator / (denominator + K.epsilon())\n</code></pre>\n\n<p>```</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 496892,
          "author_name": "Andrii Sydorchuk",
          "author_url": "",
          "post_date": "2019-03-22T17:58:54.890000",
          "content": "<p>Jack, I am not sure ROC AUC would have been better, since it would require you to pick a threshold and a few samples on the wrong side of the threshold would change result significantly. So it has essentially the same downside as MCC when you have so few positive samples.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 496895,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-03-22T18:06:52.370000",
          "content": "<p>I think maybe clipping strategy may play a big role on the sensitivity of prediction results: <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/48241\">https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/48241</a></p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 496936,
          "author_name": "Andrii Sydorchuk",
          "author_url": "",
          "post_date": "2019-03-22T18:52:01.337000",
          "content": "<p>Correct, the point was that competition results would be more stable if continuous metric was used for evaluation (e.g. RMSE, logloss, etc).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 496988,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-03-22T19:59:31.793000",
          "content": "<p><a href=\"/asydorchuk\">@asydorchuk</a> what if it were like ROC AUC in the Santander competition were you  submit probabilities rather than 0 or 1? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 497017,
          "author_name": "Andrii Sydorchuk",
          "author_url": "",
          "post_date": "2019-03-22T21:06:31.460000",
          "content": "<p>Yes, that should improve evaluation stability.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 497165,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2019-03-23T03:54:21.627000",
          "content": "<p>I had the same thought about submitting probabilities and scoring them by ROC AUC or some other continuous metric.  Although I'm not an expert on detecting power line faults, it seems to me that rounding predicted probabilities to 0 or 1 discards a lot of potentially useful information.  I wonder what the reasoning was behind using the MCC.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 497379,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-23T13:21:26.857000",
      "content": "<p>Cool function! I will use it to predict my score in all future competitions :-)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496510": "```\nimport numpy as np\n\ndef final_standings(participants, solutions):\n    np.random.shuffle(participants)\n    return participants\n```\n\nJokes aside, I am a bit disappointed with quality of data in this competition. It seems like final standings are mainly random.\n\nNevertheless congrats to all winners!",
    "496670": "A lot of it is down to the metric rather than the data. Other Kaggle competitions that used MCC to evaluate submissions had big shake-ups too.",
    "497379": "Cool function! I will use it to predict my score in all future competitions :-)\n"
  }
}