{
  "id": 79102,
  "title": "Model instability",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/79102",
  "author_name": "",
  "post_date": "2019-01-31T00:17:58.363496Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Has anyone else experienced difficulties with the cv/lb differences? It seems that the Matthews correlation varies wildly between epochs and between cv and on the lb. For instance, I just changed the seed on the top performing public kernel and the correlation went from 0.694 to 0.671. This is worrisome. </p>",
  "messages": [
    {
      "id": "463952",
      "postDate": "01/31/2019 00:17:58",
      "content": "<p>Has anyone else experienced difficulties with the cv/lb differences? It seems that the Matthews correlation varies wildly between epochs and between cv and on the lb. For instance, I just changed the seed on the top performing public kernel and the correlation went from 0.694 to 0.671. This is worrisome. </p>",
      "rawMarkdown": "Has anyone else experienced difficulties with the cv/lb differences? It seems that the Matthews correlation varies wildly between epochs and between cv and on the lb. For instance, I just changed the seed on the top performing public kernel and the correlation went from 0.694 to 0.671. This is worrisome.",
      "votes": null
    },
    {
      "id": "463958",
      "postDate": "01/31/2019 00:31:21",
      "content": "<p>The difference between 0.671 and 0.694 is not huge in the scope of what you're describing. Also you mention differences between CV and LB, but that's pretty common even if your model is well generalizable just because of noise and such. Since you're referring to epochs I'm assuming you're using neural networks. Depending on their size, gradient descent algorithm, and architecture, they could be more or less vulnerable to converging to a local minima. In a situation where changing the random seed causes performance to change, that's what you're experiencing, convergence to different local minima. </p>\n\n<p>There's no guaranteed fix to such an issue but there are several things that can help. Adjusting your learning rate and gradient descent algorithm can help with such issues. Different neural network architectures can also be affected in different ways by a irregular non-concave error function, so changes in that regard may help you. Try looking into other alternative models or the possibility of ensemble models (stacking, boosting, bagging) to improve your generalization out-of-sample. I'd stop fiddling around with your random seed and focus more on the nature of your solution to the problem, a difference of a few hundredths of a data points due to changes in random seed will more-often-than-not be representative of luck by changing convergence location slightly, and at that point it's a coin flip as to whether your CV/LB will be helped or hurt. </p>",
      "rawMarkdown": "The difference between 0.671 and 0.694 is not huge in the scope of what you're describing. Also you mention differences between CV and LB, but that's pretty common even if your model is well generalizable just because of noise and such. Since you're referring to epochs I'm assuming you're using neural networks. Depending on their size, gradient descent algorithm, and architecture, they could be more or less vulnerable to converging to a local minima. In a situation where changing the random seed causes performance to change, that's what you're experiencing, convergence to different local minima. \n\nThere's no guaranteed fix to such an issue but there are several things that can help. Adjusting your learning rate and gradient descent algorithm can help with such issues. Different neural network architectures can also be affected in different ways by a irregular non-concave error function, so changes in that regard may help you. Try looking into other alternative models or the possibility of ensemble models (stacking, boosting, bagging) to improve your generalization out-of-sample. I'd stop fiddling around with your random seed and focus more on the nature of your solution to the problem, a difference of a few hundredths of a data points due to changes in random seed will more-often-than-not be representative of luck by changing convergence location slightly, and at that point it's a coin flip as to whether your CV/LB will be helped or hurt.",
      "votes": null
    },
    {
      "id": "463981",
      "postDate": "01/31/2019 02:06:58",
      "content": "<p>I guess less sample size is the reason. I have experienced value in range of 0.657 to 0.724. To get reproducible result i fixed the seed but didn't get more than .684 so far. </p>",
      "rawMarkdown": "I guess less sample size is the reason. I have experienced value in range of 0.657 to 0.724. To get reproducible result i fixed the seed but didn't get more than .684 so far.",
      "votes": null
    },
    {
      "id": "464264",
      "postDate": "01/31/2019 14:02:13",
      "content": "<p>Yeah, this competition is hard. I still can't get a normal pipeline, where CV an LB are correlated. :( </p>",
      "rawMarkdown": "Yeah, this competition is hard. I still can't get a normal pipeline, where CV an LB are correlated. :(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 463958,
      "author_name": "alecthekulak",
      "author_url": "",
      "post_date": "01/31/2019 00:31:21",
      "content": "<p>The difference between 0.671 and 0.694 is not huge in the scope of what you're describing. Also you mention differences between CV and LB, but that's pretty common even if your model is well generalizable just because of noise and such. Since you're referring to epochs I'm assuming you're using neural networks. Depending on their size, gradient descent algorithm, and architecture, they could be more or less vulnerable to converging to a local minima. In a situation where changing the random seed causes performance to change, that's what you're experiencing, convergence to different local minima. </p>\n\n<p>There's no guaranteed fix to such an issue but there are several things that can help. Adjusting your learning rate and gradient descent algorithm can help with such issues. Different neural network architectures can also be affected in different ways by a irregular non-concave error function, so changes in that regard may help you. Try looking into other alternative models or the possibility of ensemble models (stacking, boosting, bagging) to improve your generalization out-of-sample. I'd stop fiddling around with your random seed and focus more on the nature of your solution to the problem, a difference of a few hundredths of a data points due to changes in random seed will more-often-than-not be representative of luck by changing convergence location slightly, and at that point it's a coin flip as to whether your CV/LB will be helped or hurt. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 463981,
      "author_name": "harshit92",
      "author_url": "",
      "post_date": "01/31/2019 02:06:58",
      "content": "<p>I guess less sample size is the reason. I have experienced value in range of 0.657 to 0.724. To get reproducible result i fixed the seed but didn't get more than .684 so far. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 464264,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "01/31/2019 14:02:13",
      "content": "<p>Yeah, this competition is hard. I still can't get a normal pipeline, where CV an LB are correlated. :( </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "463952": "Has anyone else experienced difficulties with the cv/lb differences? It seems that the Matthews correlation varies wildly between epochs and between cv and on the lb. For instance, I just changed the seed on the top performing public kernel and the correlation went from 0.694 to 0.671. This is worrisome.",
    "463958": "The difference between 0.671 and 0.694 is not huge in the scope of what you're describing. Also you mention differences between CV and LB, but that's pretty common even if your model is well generalizable just because of noise and such. Since you're referring to epochs I'm assuming you're using neural networks. Depending on their size, gradient descent algorithm, and architecture, they could be more or less vulnerable to converging to a local minima. In a situation where changing the random seed causes performance to change, that's what you're experiencing, convergence to different local minima. \n\nThere's no guaranteed fix to such an issue but there are several things that can help. Adjusting your learning rate and gradient descent algorithm can help with such issues. Different neural network architectures can also be affected in different ways by a irregular non-concave error function, so changes in that regard may help you. Try looking into other alternative models or the possibility of ensemble models (stacking, boosting, bagging) to improve your generalization out-of-sample. I'd stop fiddling around with your random seed and focus more on the nature of your solution to the problem, a difference of a few hundredths of a data points due to changes in random seed will more-often-than-not be representative of luck by changing convergence location slightly, and at that point it's a coin flip as to whether your CV/LB will be helped or hurt.",
    "463981": "I guess less sample size is the reason. I have experienced value in range of 0.657 to 0.724. To get reproducible result i fixed the seed but didn't get more than .684 so far.",
    "464264": "Yeah, this competition is hard. I still can't get a normal pipeline, where CV an LB are correlated. :("
  },
  "source": "meta"
}