{
  "id": 101268,
  "title": "high score variance ",
  "url": "/competitions/aptos2019-blindness-detection/discussion/101268",
  "author_name": "",
  "post_date": "2019-07-24T14:14:36.518862400Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hello everybody,</p>\n\n<p>I just realized at this competition, that there is a high variance of different models learning the relationship of the underlying differences. I just started with keras implementations of efficientnet as a regression problem, with minor augmentation and got max. a score of 0.6. Then switching to pytorch gave a LB score of 0.788. I tried different models like ResXNet, Squeeznet and own small Residual CNN implemetations. With these, I can't even reach a kappa score of above 0.4 at a testset. Then also local training and training in kaggle kernels gives me a huge difference in score. The pytorch efficientnet-b0 converges to ~ 0.1 MSE, while with local training I get like min. 0.8 MSE. I just wonder why there is such a huge difference between architectures and frameworks. Is there a subtle detail, which makes a huge difference between the model converging or diverging ? </p>",
  "messages": [
    {
      "id": "583469",
      "postDate": "07/24/2019 14:14:36",
      "content": "<p>Hello everybody,</p>\n\n<p>I just realized at this competition, that there is a high variance of different models learning the relationship of the underlying differences. I just started with keras implementations of efficientnet as a regression problem, with minor augmentation and got max. a score of 0.6. Then switching to pytorch gave a LB score of 0.788. I tried different models like ResXNet, Squeeznet and own small Residual CNN implemetations. With these, I can't even reach a kappa score of above 0.4 at a testset. Then also local training and training in kaggle kernels gives me a huge difference in score. The pytorch efficientnet-b0 converges to ~ 0.1 MSE, while with local training I get like min. 0.8 MSE. I just wonder why there is such a huge difference between architectures and frameworks. Is there a subtle detail, which makes a huge difference between the model converging or diverging ? </p>",
      "rawMarkdown": "Hello everybody,\n\nI just realized at this competition, that there is a high variance of different models learning the relationship of the underlying differences. I just started with keras implementations of efficientnet as a regression problem, with minor augmentation and got max. a score of 0.6. Then switching to pytorch gave a LB score of 0.788. I tried different models like ResXNet, Squeeznet and own small Residual CNN implemetations. With these, I can't even reach a kappa score of above 0.4 at a testset. Then also local training and training in kaggle kernels gives me a huge difference in score. The pytorch efficientnet-b0 converges to ~ 0.1 MSE, while with local training I get like min. 0.8 MSE. I just wonder why there is such a huge difference between architectures and frameworks. Is there a subtle detail, which makes a huge difference between the model converging or diverging ?",
      "votes": null
    },
    {
      "id": "585281",
      "postDate": "07/27/2019 07:45:56",
      "content": "<p><a href=\"/stanislavmalorodov\">@stanislavmalorodov</a> , did you manage to find answers to your above questions. can you share them.</p>",
      "rawMarkdown": "stanislavmalorodov , did you manage to find answers to your above questions. can you share them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 585281,
      "author_name": "ravivadapalli",
      "author_url": "",
      "post_date": "07/27/2019 07:45:56",
      "content": "<p><a href=\"/stanislavmalorodov\">@stanislavmalorodov</a> , did you manage to find answers to your above questions. can you share them.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "583469": "Hello everybody,\n\nI just realized at this competition, that there is a high variance of different models learning the relationship of the underlying differences. I just started with keras implementations of efficientnet as a regression problem, with minor augmentation and got max. a score of 0.6. Then switching to pytorch gave a LB score of 0.788. I tried different models like ResXNet, Squeeznet and own small Residual CNN implemetations. With these, I can't even reach a kappa score of above 0.4 at a testset. Then also local training and training in kaggle kernels gives me a huge difference in score. The pytorch efficientnet-b0 converges to ~ 0.1 MSE, while with local training I get like min. 0.8 MSE. I just wonder why there is such a huge difference between architectures and frameworks. Is there a subtle detail, which makes a huge difference between the model converging or diverging ?",
    "585281": "stanislavmalorodov , did you manage to find answers to your above questions. can you share them."
  },
  "source": "meta"
}