{
  "id": 106869,
  "title": "QWK of 0.97 on validation but 0.030 on submission?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106869",
  "author_name": "",
  "post_date": "2019-08-31T11:43:05.412433400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Is the test data so wildly different to the training data that a model with a QWK of 0.90+ on a holdout set of ~2k images from the 2019 dataset should score anything ranging from -0.017 to 0.030 when submitted to the competition or should I be looking for some kind of submission issue in the code instead?</p>\n\n<p>This is a classification DenseNet-121 in Keras trained on the 2015 dataset and then fine tuned on the 2019 dataset, with both sets being balanced using image augmentation for training. I'm using the same Metrics callback that many others are using and a custom loss function that does a weighted categorical_crossentropy for ordinal classes.</p>\n\n<p>I've shared a recent version of the kernel but it's quite messy because as i've gotten more desperate i've pulled them apart to an extent where they barely resemble where I started and are now stripped down for a faster train/test/submit turnaround while I try different things.  I have separate training and inference kernels to speed up submission.</p>\n\n<p><a href=\"https://www.kaggle.com/skslater/aptos-2-classification?scriptVersionId=19885459\">Training</a>\n<a href=\"https://www.kaggle.com/skslater/kernel2c9b3e693a\">Inference</a></p>\n\n<p>On top of what's in there now i've also tried:\n- Early stopping and saving best weights for each stage based on validation loss\n- DenseNet121 vs ResNet50 (most versions have been ResNet)\n- More/less dense and dropout layers in classification\n- Ben's cropping\n- Grayscale vs RGB\n- Train on 2015, validate on 2019 (the weird two stage 2019 training you see in the current kernel is there for speed so I could run through early epochs without the validation overhead)</p>\n\n<p>I'm now 30+ versions across two training kernels and still none the wiser where i'm going wrong, especially as they seem almost identical to others that are scoring 0.70 upwards.</p>\n\n<p>Any pointers would be much appreciated as i'm fairly new to this and as much as i'd love to figure it out myself i'm starting to feel like i've missed something fundamental.</p>",
  "messages": [
    {
      "id": "614347",
      "postDate": "08/31/2019 11:43:05",
      "content": "<p>Is the test data so wildly different to the training data that a model with a QWK of 0.90+ on a holdout set of ~2k images from the 2019 dataset should score anything ranging from -0.017 to 0.030 when submitted to the competition or should I be looking for some kind of submission issue in the code instead?</p>\n\n<p>This is a classification DenseNet-121 in Keras trained on the 2015 dataset and then fine tuned on the 2019 dataset, with both sets being balanced using image augmentation for training. I'm using the same Metrics callback that many others are using and a custom loss function that does a weighted categorical_crossentropy for ordinal classes.</p>\n\n<p>I've shared a recent version of the kernel but it's quite messy because as i've gotten more desperate i've pulled them apart to an extent where they barely resemble where I started and are now stripped down for a faster train/test/submit turnaround while I try different things.  I have separate training and inference kernels to speed up submission.</p>\n\n<p><a href=\"https://www.kaggle.com/skslater/aptos-2-classification?scriptVersionId=19885459\">Training</a>\n<a href=\"https://www.kaggle.com/skslater/kernel2c9b3e693a\">Inference</a></p>\n\n<p>On top of what's in there now i've also tried:\n- Early stopping and saving best weights for each stage based on validation loss\n- DenseNet121 vs ResNet50 (most versions have been ResNet)\n- More/less dense and dropout layers in classification\n- Ben's cropping\n- Grayscale vs RGB\n- Train on 2015, validate on 2019 (the weird two stage 2019 training you see in the current kernel is there for speed so I could run through early epochs without the validation overhead)</p>\n\n<p>I'm now 30+ versions across two training kernels and still none the wiser where i'm going wrong, especially as they seem almost identical to others that are scoring 0.70 upwards.</p>\n\n<p>Any pointers would be much appreciated as i'm fairly new to this and as much as i'd love to figure it out myself i'm starting to feel like i've missed something fundamental.</p>",
      "rawMarkdown": "Is the test data so wildly different to the training data that a model with a QWK of 0.90+ on a holdout set of ~2k images from the 2019 dataset should score anything ranging from -0.017 to 0.030 when submitted to the competition or should I be looking for some kind of submission issue in the code instead?\n\nThis is a classification DenseNet-121 in Keras trained on the 2015 dataset and then fine tuned on the 2019 dataset, with both sets being balanced using image augmentation for training. I'm using the same Metrics callback that many others are using and a custom loss function that does a weighted categorical_crossentropy for ordinal classes.\n\nI've shared a recent version of the kernel but it's quite messy because as i've gotten more desperate i've pulled them apart to an extent where they barely resemble where I started and are now stripped down for a faster train/test/submit turnaround while I try different things.  I have separate training and inference kernels to speed up submission.\n\n[Training](https://www.kaggle.com/skslater/aptos-2-classification?scriptVersionId=19885459)\n[Inference](https://www.kaggle.com/skslater/kernel2c9b3e693a)\n\nOn top of what's in there now i've also tried:\n- Early stopping and saving best weights for each stage based on validation loss\n- DenseNet121 vs ResNet50 (most versions have been ResNet)\n- More/less dense and dropout layers in classification\n- Ben's cropping\n- Grayscale vs RGB\n- Train on 2015, validate on 2019 (the weird two stage 2019 training you see in the current kernel is there for speed so I could run through early epochs without the validation overhead)\n\nI'm now 30+ versions across two training kernels and still none the wiser where i'm going wrong, especially as they seem almost identical to others that are scoring 0.70 upwards.\n\nAny pointers would be much appreciated as i'm fairly new to this and as much as i'd love to figure it out myself i'm starting to feel like i've missed something fundamental.",
      "votes": null
    },
    {
      "id": "614398",
      "postDate": "08/31/2019 12:41:29",
      "content": "<p>refer this post: <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-608081\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-608081</a>. May be you will get something here.</p>",
      "rawMarkdown": "refer this post: https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-608081. May be you will get something here.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 614398,
      "author_name": "suchith0312",
      "author_url": "",
      "post_date": "08/31/2019 12:41:29",
      "content": "<p>refer this post: <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-608081\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-608081</a>. May be you will get something here.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "614347": "Is the test data so wildly different to the training data that a model with a QWK of 0.90+ on a holdout set of ~2k images from the 2019 dataset should score anything ranging from -0.017 to 0.030 when submitted to the competition or should I be looking for some kind of submission issue in the code instead?\n\nThis is a classification DenseNet-121 in Keras trained on the 2015 dataset and then fine tuned on the 2019 dataset, with both sets being balanced using image augmentation for training. I'm using the same Metrics callback that many others are using and a custom loss function that does a weighted categorical_crossentropy for ordinal classes.\n\nI've shared a recent version of the kernel but it's quite messy because as i've gotten more desperate i've pulled them apart to an extent where they barely resemble where I started and are now stripped down for a faster train/test/submit turnaround while I try different things.  I have separate training and inference kernels to speed up submission.\n\n[Training](https://www.kaggle.com/skslater/aptos-2-classification?scriptVersionId=19885459)\n[Inference](https://www.kaggle.com/skslater/kernel2c9b3e693a)\n\nOn top of what's in there now i've also tried:\n- Early stopping and saving best weights for each stage based on validation loss\n- DenseNet121 vs ResNet50 (most versions have been ResNet)\n- More/less dense and dropout layers in classification\n- Ben's cropping\n- Grayscale vs RGB\n- Train on 2015, validate on 2019 (the weird two stage 2019 training you see in the current kernel is there for speed so I could run through early epochs without the validation overhead)\n\nI'm now 30+ versions across two training kernels and still none the wiser where i'm going wrong, especially as they seem almost identical to others that are scoring 0.70 upwards.\n\nAny pointers would be much appreciated as i'm fairly new to this and as much as i'd love to figure it out myself i'm starting to feel like i've missed something fundamental.",
    "614398": "refer this post: https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100815#latest-608081. May be you will get something here."
  },
  "source": "meta"
}