{
  "id": 106900,
  "title": "More data but well worse results on same model",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106900",
  "author_name": "",
  "post_date": "2019-08-31T17:00:53.313581Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I took a notebook that gave me a 0.733 result with the current competition data and used the (much larger ) dataset from the previous diabetic blindness competition instead (I didn't merge them) and I got results that were at best 0.5. The model is unchanged but the learning is just awful. I am not even talking about the submission score. I am talking about the validation and quadratic scores when training it.  In some cases, the quadratics were in the .2 -.3 range after 7 epochs and a previously good learning rate. <br>\nI can't figure out why. Shouldn't more data give me a better result not far worse ?</p>",
  "messages": [
    {
      "id": "614542",
      "postDate": "08/31/2019 17:00:53",
      "content": "<p>I took a notebook that gave me a 0.733 result with the current competition data and used the (much larger ) dataset from the previous diabetic blindness competition instead (I didn't merge them) and I got results that were at best 0.5. The model is unchanged but the learning is just awful. I am not even talking about the submission score. I am talking about the validation and quadratic scores when training it.  In some cases, the quadratics were in the .2 -.3 range after 7 epochs and a previously good learning rate. <br>\nI can't figure out why. Shouldn't more data give me a better result not far worse ?</p>",
      "rawMarkdown": "I took a notebook that gave me a 0.733 result with the current competition data and used the (much larger ) dataset from the previous diabetic blindness competition instead (I didn't merge them) and I got results that were at best 0.5. The model is unchanged but the learning is just awful. I am not even talking about the submission score. I am talking about the validation and quadratic scores when training it.  In some cases, the quadratics were in the .2 -.3 range after 7 epochs and a previously good learning rate.  \nI can't figure out why. Shouldn't more data give me a better result not far worse ?",
      "votes": null
    },
    {
      "id": "614574",
      "postDate": "08/31/2019 17:39:05",
      "content": "<p>A huge part of this competition is normalizing (scaling, colorizing, zooming, ect.) images. Our team has spent weeks tweaking our image processing and comparing the test data.</p>",
      "rawMarkdown": "A huge part of this competition is normalizing (scaling, colorizing, zooming, ect.) images. Our team has spent weeks tweaking our image processing and comparing the test data.",
      "votes": null
    },
    {
      "id": "614763",
      "postDate": "09/01/2019 03:12:03",
      "content": "<p>I am facing the same issue. On incrementally adding 2015 data to 2019 data, accuracy reduces for me. Will let you know if something works.</p>",
      "rawMarkdown": "I am facing the same issue. On incrementally adding 2015 data to 2019 data, accuracy reduces for me. Will let you know if something works.",
      "votes": null
    },
    {
      "id": "614852",
      "postDate": "09/01/2019 06:47:36",
      "content": "<p>I think fast.ai does at least some of that. What deep learning library are you using? </p>",
      "rawMarkdown": "I think fast.ai does at least some of that. What deep learning library are you using?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 614574,
      "author_name": "ticksterlee",
      "author_url": "",
      "post_date": "08/31/2019 17:39:05",
      "content": "<p>A huge part of this competition is normalizing (scaling, colorizing, zooming, ect.) images. Our team has spent weeks tweaking our image processing and comparing the test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 614852,
          "author_name": "chrisfs",
          "author_url": "",
          "post_date": "09/01/2019 06:47:36",
          "content": "<p>I think fast.ai does at least some of that. What deep learning library are you using? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614763,
      "author_name": "anuragtr",
      "author_url": "",
      "post_date": "09/01/2019 03:12:03",
      "content": "<p>I am facing the same issue. On incrementally adding 2015 data to 2019 data, accuracy reduces for me. Will let you know if something works.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "614542": "I took a notebook that gave me a 0.733 result with the current competition data and used the (much larger ) dataset from the previous diabetic blindness competition instead (I didn't merge them) and I got results that were at best 0.5. The model is unchanged but the learning is just awful. I am not even talking about the submission score. I am talking about the validation and quadratic scores when training it.  In some cases, the quadratics were in the .2 -.3 range after 7 epochs and a previously good learning rate.  \nI can't figure out why. Shouldn't more data give me a better result not far worse ?",
    "614574": "A huge part of this competition is normalizing (scaling, colorizing, zooming, ect.) images. Our team has spent weeks tweaking our image processing and comparing the test data.",
    "614763": "I am facing the same issue. On incrementally adding 2015 data to 2019 data, accuracy reduces for me. Will let you know if something works.",
    "614852": "I think fast.ai does at least some of that. What deep learning library are you using?"
  },
  "source": "meta"
}