{
  "id": 78562,
  "title": "Loss vs MCC on Neural Networks",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/78562",
  "author_name": "",
  "post_date": "2019-01-25T11:57:48.537510800Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am experimenting with Neural Networks (mostly LSTM and CNN) to tackle this competition and I have noticed some weird (?) behavior on the relation loss (binary cross entropy)/metric (MCC).</p>\n\n<p>What happens is that although the loss has a downwards trends in general, the metric has a kind of random behavior within each epoch of training. This means that sometimes the metric has a better score even with higher values of loss. This difference can be quite high in some cases. I have also noticed the same behavior on some kernels posted in this competition.</p>\n\n<p>I understand that smaller loss doesn't necessary translates to a better metric, but my past experience thought me that the trend tends to follow, meaning that if the loss gets low enough, the metric tends to follow. The problem here is to decide when to save the weights of the network, with the best validation loss or the best validation metric?</p>\n\n<p>Obviously saving the weights on the best validation metric gives a higher CV result, but I have a feeling this might be leading to overfitting. Anyone have any thoughts on this matter?</p>\n\n<p>Also, I'm using CPMP amazing code to <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682\">optimize the probabilites for best MCC</a>, not sure if this could affect this somehow.</p>",
  "messages": [
    {
      "id": "461156",
      "postDate": "01/25/2019 11:57:48",
      "content": "<p>I am experimenting with Neural Networks (mostly LSTM and CNN) to tackle this competition and I have noticed some weird (?) behavior on the relation loss (binary cross entropy)/metric (MCC).</p>\n\n<p>What happens is that although the loss has a downwards trends in general, the metric has a kind of random behavior within each epoch of training. This means that sometimes the metric has a better score even with higher values of loss. This difference can be quite high in some cases. I have also noticed the same behavior on some kernels posted in this competition.</p>\n\n<p>I understand that smaller loss doesn't necessary translates to a better metric, but my past experience thought me that the trend tends to follow, meaning that if the loss gets low enough, the metric tends to follow. The problem here is to decide when to save the weights of the network, with the best validation loss or the best validation metric?</p>\n\n<p>Obviously saving the weights on the best validation metric gives a higher CV result, but I have a feeling this might be leading to overfitting. Anyone have any thoughts on this matter?</p>\n\n<p>Also, I'm using CPMP amazing code to <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682\">optimize the probabilites for best MCC</a>, not sure if this could affect this somehow.</p>",
      "rawMarkdown": "I am experimenting with Neural Networks (mostly LSTM and CNN) to tackle this competition and I have noticed some weird (?) behavior on the relation loss (binary cross entropy)/metric (MCC).\n\nWhat happens is that although the loss has a downwards trends in general, the metric has a kind of random behavior within each epoch of training. This means that sometimes the metric has a better score even with higher values of loss. This difference can be quite high in some cases. I have also noticed the same behavior on some kernels posted in this competition.\n\nI understand that smaller loss doesn't necessary translates to a better metric, but my past experience thought me that the trend tends to follow, meaning that if the loss gets low enough, the metric tends to follow. The problem here is to decide when to save the weights of the network, with the best validation loss or the best validation metric?\n\nObviously saving the weights on the best validation metric gives a higher CV result, but I have a feeling this might be leading to overfitting. Anyone have any thoughts on this matter?\n\nAlso, I'm using CPMP amazing code to [optimize the probabilites for best MCC][1], not sure if this could affect this somehow.\n\n\n  [1]: https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682",
      "votes": null
    },
    {
      "id": "466205",
      "postDate": "02/04/2019 21:22:11",
      "content": "<p>Hi I wrote a article about my new model for CNN, Inside that article I found the same thing and explained it as butterfly effect here is the link, <a href=\"http://vixra.org/abs/1901.0463?ref=10380359\">http://vixra.org/abs/1901.0463?ref=10380359</a></p>",
      "rawMarkdown": "Hi I wrote a article about my new model for CNN, Inside that article I found the same thing and explained it as butterfly effect here is the link, http://vixra.org/abs/1901.0463?ref=10380359",
      "votes": null
    },
    {
      "id": "466206",
      "postDate": "02/04/2019 21:25:48",
      "content": "<p>Maybe there are a little different, the difference is that the divergence in my experiment is huge , but I think , the reason should be same, no linear system have this butterfly effect and reverse one </p>",
      "rawMarkdown": "Maybe there are a little different, the difference is that the divergence in my experiment is huge , but I think , the reason should be same, no linear system have this butterfly effect and reverse one",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 466205,
      "author_name": "dotkingtech",
      "author_url": "",
      "post_date": "02/04/2019 21:22:11",
      "content": "<p>Hi I wrote a article about my new model for CNN, Inside that article I found the same thing and explained it as butterfly effect here is the link, <a href=\"http://vixra.org/abs/1901.0463?ref=10380359\">http://vixra.org/abs/1901.0463?ref=10380359</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 466206,
      "author_name": "dotkingtech",
      "author_url": "",
      "post_date": "02/04/2019 21:25:48",
      "content": "<p>Maybe there are a little different, the difference is that the divergence in my experiment is huge , but I think , the reason should be same, no linear system have this butterfly effect and reverse one </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "461156": "I am experimenting with Neural Networks (mostly LSTM and CNN) to tackle this competition and I have noticed some weird (?) behavior on the relation loss (binary cross entropy)/metric (MCC).\n\nWhat happens is that although the loss has a downwards trends in general, the metric has a kind of random behavior within each epoch of training. This means that sometimes the metric has a better score even with higher values of loss. This difference can be quite high in some cases. I have also noticed the same behavior on some kernels posted in this competition.\n\nI understand that smaller loss doesn't necessary translates to a better metric, but my past experience thought me that the trend tends to follow, meaning that if the loss gets low enough, the metric tends to follow. The problem here is to decide when to save the weights of the network, with the best validation loss or the best validation metric?\n\nObviously saving the weights on the best validation metric gives a higher CV result, but I have a feeling this might be leading to overfitting. Anyone have any thoughts on this matter?\n\nAlso, I'm using CPMP amazing code to [optimize the probabilites for best MCC][1], not sure if this could affect this somehow.\n\n\n  [1]: https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76682",
    "466205": "Hi I wrote a article about my new model for CNN, Inside that article I found the same thing and explained it as butterfly effect here is the link, http://vixra.org/abs/1901.0463?ref=10380359",
    "466206": "Maybe there are a little different, the difference is that the divergence in my experiment is huge , but I think , the reason should be same, no linear system have this butterfly effect and reverse one"
  },
  "source": "meta"
}