{
  "id": 716684,
  "title": " Need help with a 5-class Diabetic Retinopathy model – Mixed predictions despite trying multiple fixes",
  "url": "/competitions/aptos2019-blindness-detection/discussion/716684",
  "author_name": "Coven Chaos",
  "post_date": "2026-06-30T20:50:00.174000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm a final-year Computer Engineering student building a Flask-based AI Diabetic Retinopathy Detection system. The web application itself is complete with patient management, authentication, dashboard, PDF report generation, prediction history, and AI inference.</p>\n<p>The only issue I'm facing is with the AI model.</p>\n<p>I'm using a 5-class Diabetic Retinopathy classifier trained on the APTOS 2019 dataset.</p>\n<p>Classes:</p>\n<p>No DR</p>\n<p>Mild</p>\n<p>Moderate</p>\n<p>Severe</p>\n<p>Proliferative DR</p>\n<p>The model predicts all five classes, but the predictions are inconsistent.</p>\n<p>Examples:</p>\n<p>Moderate is sometimes classified as Severe or Proliferative.</p>\n<p>Severe is often classified as Moderate or Proliferative and is rarely predicted correctly.</p>\n<p>Some fundus images from outside the APTOS dataset produce completely unexpected results.</p>\n<p>The model sometimes shows very high confidence (90%+) even when the prediction appears incorrect.</p>\n<p>Things I've already tried:</p>\n<p>Different pretrained models (including a ResNet50 trained on APTOS)</p>\n<p>ResNet152 implementation</p>\n<p>Correct preprocessing (RGB conversion, resizing, normalization)</p>\n<p>Verified class mapping</p>\n<p>Softmax confidence scores</p>\n<p>Test-Time Augmentation (TTA)</p>\n<p>Image quality validation</p>\n<p>Top-3 predictions instead of only one prediction</p>\n<p>I'm trying to understand whether this is:</p>\n<p>A domain shift problem between APTOS and other datasets?</p>\n<p>A limitation of the pretrained model?</p>\n<p>A preprocessing issue?</p>\n<p>Class imbalance?</p>\n<p>Or simply expected behavior in 5-class DR classification?</p>\n<p>I'm also considering using an ensemble (ResNet50 + EfficientNet + DenseNet), but it's difficult to find compatible pretrained 5-class diabetic retinopathy models.</p>\n<p>I'd really appreciate advice from anyone who has worked on retinal image classification or medical AI.</p>\n<p>My questions are:</p>\n<ol>\n<li><p>Is this level of class confusion common in diabetic retinopathy models?</p></li>\n<li><p>What preprocessing techniques made the biggest improvement for you (CLAHE, retinal cropping, illumination correction, etc.)?</p></li>\n<li><p>Has anyone significantly improved results using ensemble models?</p></li>\n<li><p>Are there any high-quality pretrained 5-class DR models that you'd recommend?</p></li>\n<li><p>If you were in my situation, what would be the first thing you'd investigate to improve prediction consistency?</p></li>\n</ol>\n<p>Any suggestions, GitHub repositories, pretrained models, research papers, or personal experiences would be greatly appreciated.</p>\n<p>Thanks </p>",
  "messages": [
    {
      "id": 3484891,
      "postDate": "2026-06-30T20:50:00.173Z",
      "content": "<p>Hi everyone,</p>\n<p>I'm a final-year Computer Engineering student building a Flask-based AI Diabetic Retinopathy Detection system. The web application itself is complete with patient management, authentication, dashboard, PDF report generation, prediction history, and AI inference.</p>\n<p>The only issue I'm facing is with the AI model.</p>\n<p>I'm using a 5-class Diabetic Retinopathy classifier trained on the APTOS 2019 dataset.</p>\n<p>Classes:</p>\n<p>No DR</p>\n<p>Mild</p>\n<p>Moderate</p>\n<p>Severe</p>\n<p>Proliferative DR</p>\n<p>The model predicts all five classes, but the predictions are inconsistent.</p>\n<p>Examples:</p>\n<p>Moderate is sometimes classified as Severe or Proliferative.</p>\n<p>Severe is often classified as Moderate or Proliferative and is rarely predicted correctly.</p>\n<p>Some fundus images from outside the APTOS dataset produce completely unexpected results.</p>\n<p>The model sometimes shows very high confidence (90%+) even when the prediction appears incorrect.</p>\n<p>Things I've already tried:</p>\n<p>Different pretrained models (including a ResNet50 trained on APTOS)</p>\n<p>ResNet152 implementation</p>\n<p>Correct preprocessing (RGB conversion, resizing, normalization)</p>\n<p>Verified class mapping</p>\n<p>Softmax confidence scores</p>\n<p>Test-Time Augmentation (TTA)</p>\n<p>Image quality validation</p>\n<p>Top-3 predictions instead of only one prediction</p>\n<p>I'm trying to understand whether this is:</p>\n<p>A domain shift problem between APTOS and other datasets?</p>\n<p>A limitation of the pretrained model?</p>\n<p>A preprocessing issue?</p>\n<p>Class imbalance?</p>\n<p>Or simply expected behavior in 5-class DR classification?</p>\n<p>I'm also considering using an ensemble (ResNet50 + EfficientNet + DenseNet), but it's difficult to find compatible pretrained 5-class diabetic retinopathy models.</p>\n<p>I'd really appreciate advice from anyone who has worked on retinal image classification or medical AI.</p>\n<p>My questions are:</p>\n<ol>\n<li><p>Is this level of class confusion common in diabetic retinopathy models?</p></li>\n<li><p>What preprocessing techniques made the biggest improvement for you (CLAHE, retinal cropping, illumination correction, etc.)?</p></li>\n<li><p>Has anyone significantly improved results using ensemble models?</p></li>\n<li><p>Are there any high-quality pretrained 5-class DR models that you'd recommend?</p></li>\n<li><p>If you were in my situation, what would be the first thing you'd investigate to improve prediction consistency?</p></li>\n</ol>\n<p>Any suggestions, GitHub repositories, pretrained models, research papers, or personal experiences would be greatly appreciated.</p>\n<p>Thanks </p>",
      "rawMarkdown": "Hi everyone,\n\nI'm a final-year Computer Engineering student building a Flask-based AI Diabetic Retinopathy Detection system. The web application itself is complete with patient management, authentication, dashboard, PDF report generation, prediction history, and AI inference.\n\nThe only issue I'm facing is with the AI model.\n\nI'm using a 5-class Diabetic Retinopathy classifier trained on the APTOS 2019 dataset.\n\nClasses:\n\nNo DR\n\nMild\n\nModerate\n\nSevere\n\nProliferative DR\n\n\nThe model predicts all five classes, but the predictions are inconsistent.\n\nExamples:\n\nModerate is sometimes classified as Severe or Proliferative.\n\nSevere is often classified as Moderate or Proliferative and is rarely predicted correctly.\n\nSome fundus images from outside the APTOS dataset produce completely unexpected results.\n\nThe model sometimes shows very high confidence (90%+) even when the prediction appears incorrect.\n\n\nThings I've already tried:\n\nDifferent pretrained models (including a ResNet50 trained on APTOS)\n\nResNet152 implementation\n\nCorrect preprocessing (RGB conversion, resizing, normalization)\n\nVerified class mapping\n\nSoftmax confidence scores\n\nTest-Time Augmentation (TTA)\n\nImage quality validation\n\nTop-3 predictions instead of only one prediction\n\n\nI'm trying to understand whether this is:\n\nA domain shift problem between APTOS and other datasets?\n\nA limitation of the pretrained model?\n\nA preprocessing issue?\n\nClass imbalance?\n\nOr simply expected behavior in 5-class DR classification?\n\n\nI'm also considering using an ensemble (ResNet50 + EfficientNet + DenseNet), but it's difficult to find compatible pretrained 5-class diabetic retinopathy models.\n\nI'd really appreciate advice from anyone who has worked on retinal image classification or medical AI.\n\nMy questions are:\n\n1. Is this level of class confusion common in diabetic retinopathy models?\n\n\n2. What preprocessing techniques made the biggest improvement for you (CLAHE, retinal cropping, illumination correction, etc.)?\n\n\n3. Has anyone significantly improved results using ensemble models?\n\n\n4. Are there any high-quality pretrained 5-class DR models that you'd recommend?\n\n\n5. If you were in my situation, what would be the first thing you'd investigate to improve prediction consistency?\n\n\n\nAny suggestions, GitHub repositories, pretrained models, research papers, or personal experiences would be greatly appreciated.\n\nThanks "
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3484891": "Hi everyone,\n\nI'm a final-year Computer Engineering student building a Flask-based AI Diabetic Retinopathy Detection system. The web application itself is complete with patient management, authentication, dashboard, PDF report generation, prediction history, and AI inference.\n\nThe only issue I'm facing is with the AI model.\n\nI'm using a 5-class Diabetic Retinopathy classifier trained on the APTOS 2019 dataset.\n\nClasses:\n\nNo DR\n\nMild\n\nModerate\n\nSevere\n\nProliferative DR\n\n\nThe model predicts all five classes, but the predictions are inconsistent.\n\nExamples:\n\nModerate is sometimes classified as Severe or Proliferative.\n\nSevere is often classified as Moderate or Proliferative and is rarely predicted correctly.\n\nSome fundus images from outside the APTOS dataset produce completely unexpected results.\n\nThe model sometimes shows very high confidence (90%+) even when the prediction appears incorrect.\n\n\nThings I've already tried:\n\nDifferent pretrained models (including a ResNet50 trained on APTOS)\n\nResNet152 implementation\n\nCorrect preprocessing (RGB conversion, resizing, normalization)\n\nVerified class mapping\n\nSoftmax confidence scores\n\nTest-Time Augmentation (TTA)\n\nImage quality validation\n\nTop-3 predictions instead of only one prediction\n\n\nI'm trying to understand whether this is:\n\nA domain shift problem between APTOS and other datasets?\n\nA limitation of the pretrained model?\n\nA preprocessing issue?\n\nClass imbalance?\n\nOr simply expected behavior in 5-class DR classification?\n\n\nI'm also considering using an ensemble (ResNet50 + EfficientNet + DenseNet), but it's difficult to find compatible pretrained 5-class diabetic retinopathy models.\n\nI'd really appreciate advice from anyone who has worked on retinal image classification or medical AI.\n\nMy questions are:\n\n1. Is this level of class confusion common in diabetic retinopathy models?\n\n\n2. What preprocessing techniques made the biggest improvement for you (CLAHE, retinal cropping, illumination correction, etc.)?\n\n\n3. Has anyone significantly improved results using ensemble models?\n\n\n4. Are there any high-quality pretrained 5-class DR models that you'd recommend?\n\n\n5. If you were in my situation, what would be the first thing you'd investigate to improve prediction consistency?\n\n\n\nAny suggestions, GitHub repositories, pretrained models, research papers, or personal experiences would be greatly appreciated.\n\nThanks "
  }
}