{
  "id": 104291,
  "title": "Surprising Kappa score on the test set",
  "url": "/competitions/aptos2019-blindness-detection/discussion/104291",
  "author_name": "Arnaud S ",
  "post_date": "2019-08-15T20:28:25.626000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm designing a CNN model and fitting it on a training set  (30M images) from a data generator.\nVariance of the model is quite low (94% on train set and 90% on test set).\nThen I'm running the model for predictions on the APTOS training dataset  (3662 images) and \nthe accuracy score is pretty suitable with 93% and a corresponding Kappa score of about 0.84.</p>\n\n<p>But finally I'm getting a Kappa score of 0.13 on the test set after submitting my model !</p>\n\n<p>How is it possible to get such a gap between these two Kappa score ?\nIs there a problem of mismatching data between training set and test set ?\nAre they coming from the same distribution or not ? </p>\n\n<p>Could anyone explain to me why such a gap ?</p>\n\n<p>Thanks a lots. </p>",
  "messages": [
    {
      "id": 600512,
      "postDate": "2019-08-16T07:52:55.127Z",
      "content": "<p>Check out <a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on\">this amazing kernel</a>. May be you can get some hint.</p>\n\n<p><strong>TL;DR</strong> Your preprocessing matters a lot. There is high correlation between image size, pixel counts and train images, which does not hold for test images. So your CNN can overfit on these features.</p>\n\n<p>IMO this can be one reason.</p>",
      "rawMarkdown": "Check out [this amazing kernel](https://www.kaggle.com/taindow/be-careful-what-you-train-on). May be you can get some hint.\n\n**TL;DR** Your preprocessing matters a lot. There is high correlation between image size, pixel counts and train images, which does not hold for test images. So your CNN can overfit on these features.\n\nIMO this can be one reason.",
      "votes": 2
    },
    {
      "id": 600235,
      "postDate": "2019-08-15T20:28:25.627Z",
      "content": "<p>I'm designing a CNN model and fitting it on a training set  (30M images) from a data generator.\nVariance of the model is quite low (94% on train set and 90% on test set).\nThen I'm running the model for predictions on the APTOS training dataset  (3662 images) and \nthe accuracy score is pretty suitable with 93% and a corresponding Kappa score of about 0.84.</p>\n\n<p>But finally I'm getting a Kappa score of 0.13 on the test set after submitting my model !</p>\n\n<p>How is it possible to get such a gap between these two Kappa score ?\nIs there a problem of mismatching data between training set and test set ?\nAre they coming from the same distribution or not ? </p>\n\n<p>Could anyone explain to me why such a gap ?</p>\n\n<p>Thanks a lots. </p>",
      "rawMarkdown": "I'm designing a CNN model and fitting it on a training set  (30M images) from a data generator.\nVariance of the model is quite low (94% on train set and 90% on test set).\nThen I'm running the model for predictions on the APTOS training dataset  (3662 images) and \nthe accuracy score is pretty suitable with 93% and a corresponding Kappa score of about 0.84.\n\nBut finally I'm getting a Kappa score of 0.13 on the test set after submitting my model !\n\nHow is it possible to get such a gap between these two Kappa score ?\nIs there a problem of mismatching data between training set and test set ?\nAre they coming from the same distribution or not ? \n\nCould anyone explain to me why such a gap ?\n\nThanks a lots. "
    },
    {
      "id": 600943,
      "postDate": "2019-08-16T19:28:30.200Z",
      "content": "<p>I fit model on about 30M images issued from a data generator, then I ran it on original train data (3662 images) in order to know accuracy score on theses predictions. Of course at submission, it has been evaluated on the 1928 images in the test set.  </p>",
      "rawMarkdown": "I fit model on about 30M images issued from a data generator, then I ran it on original train data (3662 images) in order to know accuracy score on theses predictions. Of course at submission, it has been evaluated on the 1928 images in the test set.  "
    },
    {
      "id": 600939,
      "postDate": "2019-08-16T19:21:53.703Z",
      "content": "<p>Thanks you Ashwani for your advice, and I'll have a look at this kernel.\nBut my preprocessing doesn't take into account these type of features (image size, pixel counts etc..) excepted a resizing.  Overfitting is so huge whereas test images aren't so different from train images after a quick viewing . </p>",
      "rawMarkdown": "Thanks you Ashwani for your advice, and I'll have a look at this kernel.\nBut my preprocessing doesn't take into account these type of features (image size, pixel counts etc..) excepted a resizing.  Overfitting is so huge whereas test images aren't so different from train images after a quick viewing . "
    },
    {
      "id": 600303,
      "postDate": "2019-08-16T00:24:09.750Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 600512,
      "author_name": "Ashwani Pandey",
      "author_url": "",
      "post_date": "2019-08-16T07:52:55.127000",
      "content": "<p>Check out <a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on\">this amazing kernel</a>. May be you can get some hint.</p>\n\n<p><strong>TL;DR</strong> Your preprocessing matters a lot. There is high correlation between image size, pixel counts and train images, which does not hold for test images. So your CNN can overfit on these features.</p>\n\n<p>IMO this can be one reason.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 600943,
      "author_name": "Arnaud S ",
      "author_url": "",
      "post_date": "2019-08-16T19:28:30.200000",
      "content": "<p>I fit model on about 30M images issued from a data generator, then I ran it on original train data (3662 images) in order to know accuracy score on theses predictions. Of course at submission, it has been evaluated on the 1928 images in the test set.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 600939,
      "author_name": "Arnaud S ",
      "author_url": "",
      "post_date": "2019-08-16T19:21:53.703000",
      "content": "<p>Thanks you Ashwani for your advice, and I'll have a look at this kernel.\nBut my preprocessing doesn't take into account these type of features (image size, pixel counts etc..) excepted a resizing.  Overfitting is so huge whereas test images aren't so different from train images after a quick viewing . </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 600303,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-16T00:24:09.750000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "600512": "Check out [this amazing kernel](https://www.kaggle.com/taindow/be-careful-what-you-train-on). May be you can get some hint.\n\n**TL;DR** Your preprocessing matters a lot. There is high correlation between image size, pixel counts and train images, which does not hold for test images. So your CNN can overfit on these features.\n\nIMO this can be one reason.",
    "600235": "I'm designing a CNN model and fitting it on a training set  (30M images) from a data generator.\nVariance of the model is quite low (94% on train set and 90% on test set).\nThen I'm running the model for predictions on the APTOS training dataset  (3662 images) and \nthe accuracy score is pretty suitable with 93% and a corresponding Kappa score of about 0.84.\n\nBut finally I'm getting a Kappa score of 0.13 on the test set after submitting my model !\n\nHow is it possible to get such a gap between these two Kappa score ?\nIs there a problem of mismatching data between training set and test set ?\nAre they coming from the same distribution or not ? \n\nCould anyone explain to me why such a gap ?\n\nThanks a lots. ",
    "600943": "I fit model on about 30M images issued from a data generator, then I ran it on original train data (3662 images) in order to know accuracy score on theses predictions. Of course at submission, it has been evaluated on the 1928 images in the test set.  ",
    "600939": "Thanks you Ashwani for your advice, and I'll have a look at this kernel.\nBut my preprocessing doesn't take into account these type of features (image size, pixel counts etc..) excepted a resizing.  Overfitting is so huge whereas test images aren't so different from train images after a quick viewing . ",
    "600303": ""
  }
}