{
  "id": 107991,
  "title": "Good riddance",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107991",
  "author_name": "Abhishek Thakur",
  "post_date": "2019-09-08T09:50:55.600000",
  "votes": 27,
  "comment_count": 19,
  "views": 0,
  "content": "<p>APTOS is finally over and it sucked big time.\n- there was no correlation between test and training data\n- no correlation between private and public test\n- images from test were quite different from training\n- a lot of people did private probing and landed gold or good ranks (obviously, they won't accept it)\n- 0.90+ QWK is very very high and very good kappa score, also unrealistic.\n- Results are useless for the organizers</p>\n\n<p>I'm glad we jumped almost 700 places and used a single model (details will come soon). Thanks to my amazing teammates: <a href=\"/konradb\">@konradb</a> and <a href=\"/aakashnain\">@aakashnain</a> </p>",
  "messages": [
    {
      "id": 621193,
      "postDate": "2019-09-08T09:50:55.600Z",
      "content": "<p>APTOS is finally over and it sucked big time.\n- there was no correlation between test and training data\n- no correlation between private and public test\n- images from test were quite different from training\n- a lot of people did private probing and landed gold or good ranks (obviously, they won't accept it)\n- 0.90+ QWK is very very high and very good kappa score, also unrealistic.\n- Results are useless for the organizers</p>\n\n<p>I'm glad we jumped almost 700 places and used a single model (details will come soon). Thanks to my amazing teammates: <a href=\"/konradb\">@konradb</a> and <a href=\"/aakashnain\">@aakashnain</a> </p>",
      "rawMarkdown": "APTOS is finally over and it sucked big time.\n- there was no correlation between test and training data\n- no correlation between private and public test\n- images from test were quite different from training\n- a lot of people did private probing and landed gold or good ranks (obviously, they won't accept it)\n- 0.90+ QWK is very very high and very good kappa score, also unrealistic.\n- Results are useless for the organizers\n\n\nI'm glad we jumped almost 700 places and used a single model (details will come soon). Thanks to my amazing teammates: @konradb and @aakashnain ",
      "votes": 27
    },
    {
      "id": 621323,
      "postDate": "2019-09-08T12:40:14.453Z",
      "content": "<p>I totally agree with you. I still couldn't understand why admin preprocessed public test set such a crazy way. What we've solved is almost nothing to do with diabetic retinopathy, just solving a silly puzzle...🙄 </p>",
      "rawMarkdown": "I totally agree with you. I still couldn't understand why admin preprocessed public test set such a crazy way. What we've solved is almost nothing to do with diabetic retinopathy, just solving a silly puzzle...🙄 ",
      "votes": 6
    },
    {
      "id": 621219,
      "postDate": "2019-09-08T10:42:21.937Z",
      "content": "<p>I dont like the metric, but with this metric it probably was necessary to split public and private differently to prevent overfitting to the distribution. Different metric and more consistent public / private would probably be better. I dont think that results are useless though, you are a bit harsh on that imho.</p>",
      "rawMarkdown": "I dont like the metric, but with this metric it probably was necessary to split public and private differently to prevent overfitting to the distribution. Different metric and more consistent public / private would probably be better. I dont think that results are useless though, you are a bit harsh on that imho.",
      "votes": 5
    },
    {
      "id": 621198,
      "postDate": "2019-09-08T10:00:19.637Z",
      "content": "<p>What is private probing can someone explain?</p>",
      "rawMarkdown": "What is private probing can someone explain?",
      "votes": 4,
      "replies": [
        {
          "id": 621752,
          "postDate": "2019-09-08T22:11:09.227Z",
          "content": "<p>Private LB probing is explained quite nicely by <a href=\"/tanlikesmath\">@tanlikesmath</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763</a></p>",
          "rawMarkdown": "Private LB probing is explained quite nicely by @tanlikesmath \n\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763"
        }
      ]
    },
    {
      "id": 621761,
      "postDate": "2019-09-08T22:54:25.553Z",
      "content": "<p>I have similar thoughts about this competition - but let's not be overly negative: you can learn something from every competition :)</p>\n\n<p>Having participated in the 2015 competition (placed 11th), in this competition the things that improved our score over time were not things that actually made a better model. \"Better model\" here I'm defining as something that I am convinced you could give any decent fundus image and it tells you the DR grade (from whatever dataset/label distribution). </p>\n\n<p>We experimented quite a lot with injecting some (light) domain knowledge into the problem. For instance, all our models in the end predicted the location of the macula and during pre-training predicted whether the eyes had advanced macular degeneration, had drusen present, or had pigment oddities/lesions. That improved our score a tiny bit, finetuning so the model fits close to the 2019 distribution helped a lot more.</p>\n\n<p>I suppose that instead of relying on our cross-validation score and intuition (and hoping the LB top end was just overfit to the weird public LB distribution) we should have followed the LB score more - turns out the distribution was the same for the hidden set. Tough luck I suppose, better luck next time :) </p>\n\n<p>If people are interested we can write up what we tried, it's a bit different from the solutions I've seen so far, and obviously it didn't score as well for this competition - achieving just under a silver medal.</p>",
      "rawMarkdown": "I have similar thoughts about this competition - but let's not be overly negative: you can learn something from every competition :)\n\nHaving participated in the 2015 competition (placed 11th), in this competition the things that improved our score over time were not things that actually made a better model. \"Better model\" here I'm defining as something that I am convinced you could give any decent fundus image and it tells you the DR grade (from whatever dataset/label distribution). \n\nWe experimented quite a lot with injecting some (light) domain knowledge into the problem. For instance, all our models in the end predicted the location of the macula and during pre-training predicted whether the eyes had advanced macular degeneration, had drusen present, or had pigment oddities/lesions. That improved our score a tiny bit, finetuning so the model fits close to the 2019 distribution helped a lot more.\n\nI suppose that instead of relying on our cross-validation score and intuition (and hoping the LB top end was just overfit to the weird public LB distribution) we should have followed the LB score more - turns out the distribution was the same for the hidden set. Tough luck I suppose, better luck next time :) \n\nIf people are interested we can write up what we tried, it's a bit different from the solutions I've seen so far, and obviously it didn't score as well for this competition - achieving just under a silver medal.",
      "votes": 1
    },
    {
      "id": 621646,
      "postDate": "2019-09-08T19:38:49.453Z",
      "content": "<p>That's true. My model that scored 0.837 on 2015 private could only get 0.74x on the 2019 public. I summarized the things I don't like about this competition in another <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107945\">thread</a>.</p>",
      "rawMarkdown": "That's true. My model that scored 0.837 on 2015 private could only get 0.74x on the 2019 public. I summarized the things I don't like about this competition in another [thread](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107945).",
      "votes": 1,
      "replies": [
        {
          "id": 621699,
          "postDate": "2019-09-08T20:49:51.400Z",
          "content": "<p>Because 2015 data is quite different from 2019 data.</p>",
          "rawMarkdown": "Because 2015 data is quite different from 2019 data.",
          "votes": 1
        },
        {
          "id": 621746,
          "postDate": "2019-09-08T21:47:38.007Z",
          "content": "<p>Also, my 2019-only trained model predicted almost all of class 1 images of 2019 as class 2. </p>",
          "rawMarkdown": "Also, my 2019-only trained model predicted almost all of class 1 images of 2019 as class 2. "
        },
        {
          "id": 621793,
          "postDate": "2019-09-09T00:36:20.390Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> By looking at all the images from the two competitions, personally I think 2015 data is of higher quality. </p>",
          "rawMarkdown": "@philippsinger By looking at all the images from the two competitions, personally I think 2015 data is of higher quality. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 621199,
      "postDate": "2019-09-08T10:02:02.750Z",
      "content": "<p>Hey, just curious, why 0.90+ QWK is unrealistic? </p>",
      "rawMarkdown": "Hey, just curious, why 0.90+ QWK is unrealistic? ",
      "votes": 1
    },
    {
      "id": 622211,
      "postDate": "2019-09-09T11:41:46.477Z",
      "content": "<p>At least my CV scores were quite correlated with private and public. I observed good correlation when I started applying Ben's preprocessing:</p>\n\n<p>Model /  CV /   Public / Private\nBlend v3     / 0.9328 / 0.813   / 0.924\nBlend v4     /  0.9368 / 0.816 / 0.927\nBlend v5      /  0.9372 / 0.817 / 0.930\nBlend v6       /  0.9379 / 0.817 / 0.930</p>",
      "rawMarkdown": "At least my CV scores were quite correlated with private and public. I observed good correlation when I started applying Ben's preprocessing:\n\nModel /  CV\t/   Public / Private\nBlend v3\t / 0.9328 / 0.813\t/ 0.924\nBlend v4\t /  0.9368 / 0.816 / 0.927\nBlend v5\t  /  0.9372 / 0.817 / 0.930\nBlend v6\t   /  0.9379 / 0.817 / 0.930",
      "votes": 2
    },
    {
      "id": 621320,
      "postDate": "2019-09-08T12:37:19.887Z",
      "content": "<p><a href=\"/abhishek\">@abhishek</a>, I share most of your sentiments. You forgot to mention the unbelievably unfair submission issues.</p>",
      "rawMarkdown": "@abhishek, I share most of your sentiments. You forgot to mention the unbelievably unfair submission issues.",
      "votes": 2
    },
    {
      "id": 621196,
      "postDate": "2019-09-08T09:56:59.663Z",
      "content": "<p>Agreed. The <code>train/test</code> distribution was a mess this time. </p>",
      "rawMarkdown": "Agreed. The `train/test` distribution was a mess this time. "
    },
    {
      "id": 621334,
      "postDate": "2019-09-08T12:46:53.770Z",
      "content": "<p>I jumped from 0.782 to 0.9... this is silly. Also I'm #1222 on the private LB but with a QWK difference wrt to the winner of barely 0.036... guess I'm satisfied with the result 👌 </p>",
      "rawMarkdown": "I jumped from 0.782 to 0.9... this is silly. Also I'm #1222 on the private LB but with a QWK difference wrt to the winner of barely 0.036... guess I'm satisfied with the result 👌 "
    },
    {
      "id": 621312,
      "postDate": "2019-09-08T12:31:18.827Z",
      "content": "<p>Hey, Abhishek thanks for your kernels. I just started competing on kaggle and your work helped me a lot.  </p>",
      "rawMarkdown": "Hey, Abhishek thanks for your kernels. I just started competing on kaggle and your work helped me a lot.  "
    },
    {
      "id": 621213,
      "postDate": "2019-09-08T10:33:03.680Z",
      "content": "<p>why results are useless for the organizers ?</p>",
      "rawMarkdown": "why results are useless for the organizers ?",
      "replies": [
        {
          "id": 621327,
          "postDate": "2019-09-08T12:42:38.797Z",
          "content": "<p>Have you verified your model with old competition data or external data (Messidor or IEEE)?\nWhen I tried, my LB score and CV score for other datasets have almost no correlation.</p>",
          "rawMarkdown": "Have you verified your model with old competition data or external data (Messidor or IEEE)?\nWhen I tried, my LB score and CV score for other datasets have almost no correlation."
        }
      ]
    },
    {
      "id": 621210,
      "postDate": "2019-09-08T10:32:19.583Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 621322,
          "postDate": "2019-09-08T12:37:45.750Z",
          "rawMarkdown": ""
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 621323,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-09-08T12:40:14.453000",
      "content": "<p>I totally agree with you. I still couldn't understand why admin preprocessed public test set such a crazy way. What we've solved is almost nothing to do with diabetic retinopathy, just solving a silly puzzle...🙄 </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 621219,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-09-08T10:42:21.937000",
      "content": "<p>I dont like the metric, but with this metric it probably was necessary to split public and private differently to prevent overfitting to the distribution. Different metric and more consistent public / private would probably be better. I dont think that results are useless though, you are a bit harsh on that imho.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 621198,
      "author_name": "Nitish Gupta",
      "author_url": "",
      "post_date": "2019-09-08T10:00:19.637000",
      "content": "<p>What is private probing can someone explain?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 621752,
          "author_name": "Sterls",
          "author_url": "",
          "post_date": "2019-09-08T22:11:09.227000",
          "content": "<p>Private LB probing is explained quite nicely by <a href=\"/tanlikesmath\">@tanlikesmath</a> </p>\n\n<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621761,
      "author_name": "Guido Zuidhof",
      "author_url": "",
      "post_date": "2019-09-08T22:54:25.553000",
      "content": "<p>I have similar thoughts about this competition - but let's not be overly negative: you can learn something from every competition :)</p>\n\n<p>Having participated in the 2015 competition (placed 11th), in this competition the things that improved our score over time were not things that actually made a better model. \"Better model\" here I'm defining as something that I am convinced you could give any decent fundus image and it tells you the DR grade (from whatever dataset/label distribution). </p>\n\n<p>We experimented quite a lot with injecting some (light) domain knowledge into the problem. For instance, all our models in the end predicted the location of the macula and during pre-training predicted whether the eyes had advanced macular degeneration, had drusen present, or had pigment oddities/lesions. That improved our score a tiny bit, finetuning so the model fits close to the 2019 distribution helped a lot more.</p>\n\n<p>I suppose that instead of relying on our cross-validation score and intuition (and hoping the LB top end was just overfit to the weird public LB distribution) we should have followed the LB score more - turns out the distribution was the same for the hidden set. Tough luck I suppose, better luck next time :) </p>\n\n<p>If people are interested we can write up what we tried, it's a bit different from the solutions I've seen so far, and obviously it didn't score as well for this competition - achieving just under a silver medal.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621646,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-09-08T19:38:49.453000",
      "content": "<p>That's true. My model that scored 0.837 on 2015 private could only get 0.74x on the 2019 public. I summarized the things I don't like about this competition in another <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107945\">thread</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 621699,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-09-08T20:49:51.400000",
          "content": "<p>Because 2015 data is quite different from 2019 data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 621746,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-09-08T21:47:38.007000",
          "content": "<p>Also, my 2019-only trained model predicted almost all of class 1 images of 2019 as class 2. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621793,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-09T00:36:20.390000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> By looking at all the images from the two competitions, personally I think 2015 data is of higher quality. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 621199,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2019-09-08T10:02:02.750000",
      "content": "<p>Hey, just curious, why 0.90+ QWK is unrealistic? </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 622211,
      "author_name": "Carlos Prades K.",
      "author_url": "",
      "post_date": "2019-09-09T11:41:46.477000",
      "content": "<p>At least my CV scores were quite correlated with private and public. I observed good correlation when I started applying Ben's preprocessing:</p>\n\n<p>Model /  CV /   Public / Private\nBlend v3     / 0.9328 / 0.813   / 0.924\nBlend v4     /  0.9368 / 0.816 / 0.927\nBlend v5      /  0.9372 / 0.817 / 0.930\nBlend v6       /  0.9379 / 0.817 / 0.930</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621320,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-09-08T12:37:19.887000",
      "content": "<p><a href=\"/abhishek\">@abhishek</a>, I share most of your sentiments. You forgot to mention the unbelievably unfair submission issues.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621196,
      "author_name": "NAIN",
      "author_url": "",
      "post_date": "2019-09-08T09:56:59.663000",
      "content": "<p>Agreed. The <code>train/test</code> distribution was a mess this time. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621334,
      "author_name": "Francesco Ramoni",
      "author_url": "",
      "post_date": "2019-09-08T12:46:53.770000",
      "content": "<p>I jumped from 0.782 to 0.9... this is silly. Also I'm #1222 on the private LB but with a QWK difference wrt to the winner of barely 0.036... guess I'm satisfied with the result 👌 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621312,
      "author_name": "Neeraj Singh Aithani",
      "author_url": "",
      "post_date": "2019-09-08T12:31:18.827000",
      "content": "<p>Hey, Abhishek thanks for your kernels. I just started competing on kaggle and your work helped me a lot.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621213,
      "author_name": "Nguyen Quan Anh Minh",
      "author_url": "",
      "post_date": "2019-09-08T10:33:03.680000",
      "content": "<p>why results are useless for the organizers ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 621327,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2019-09-08T12:42:38.797000",
          "content": "<p>Have you verified your model with old competition data or external data (Messidor or IEEE)?\nWhen I tried, my LB score and CV score for other datasets have almost no correlation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621210,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T10:32:19.583000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 621322,
          "author_name": "Neeraj Singh Aithani",
          "author_url": "",
          "post_date": "2019-09-08T12:37:45.750000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "621193": "APTOS is finally over and it sucked big time.\n- there was no correlation between test and training data\n- no correlation between private and public test\n- images from test were quite different from training\n- a lot of people did private probing and landed gold or good ranks (obviously, they won't accept it)\n- 0.90+ QWK is very very high and very good kappa score, also unrealistic.\n- Results are useless for the organizers\n\n\nI'm glad we jumped almost 700 places and used a single model (details will come soon). Thanks to my amazing teammates: @konradb and @aakashnain ",
    "621323": "I totally agree with you. I still couldn't understand why admin preprocessed public test set such a crazy way. What we've solved is almost nothing to do with diabetic retinopathy, just solving a silly puzzle...🙄 ",
    "621219": "I dont like the metric, but with this metric it probably was necessary to split public and private differently to prevent overfitting to the distribution. Different metric and more consistent public / private would probably be better. I dont think that results are useless though, you are a bit harsh on that imho.",
    "621198": "What is private probing can someone explain?",
    "621761": "I have similar thoughts about this competition - but let's not be overly negative: you can learn something from every competition :)\n\nHaving participated in the 2015 competition (placed 11th), in this competition the things that improved our score over time were not things that actually made a better model. \"Better model\" here I'm defining as something that I am convinced you could give any decent fundus image and it tells you the DR grade (from whatever dataset/label distribution). \n\nWe experimented quite a lot with injecting some (light) domain knowledge into the problem. For instance, all our models in the end predicted the location of the macula and during pre-training predicted whether the eyes had advanced macular degeneration, had drusen present, or had pigment oddities/lesions. That improved our score a tiny bit, finetuning so the model fits close to the 2019 distribution helped a lot more.\n\nI suppose that instead of relying on our cross-validation score and intuition (and hoping the LB top end was just overfit to the weird public LB distribution) we should have followed the LB score more - turns out the distribution was the same for the hidden set. Tough luck I suppose, better luck next time :) \n\nIf people are interested we can write up what we tried, it's a bit different from the solutions I've seen so far, and obviously it didn't score as well for this competition - achieving just under a silver medal.",
    "621646": "That's true. My model that scored 0.837 on 2015 private could only get 0.74x on the 2019 public. I summarized the things I don't like about this competition in another [thread](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/107945).",
    "621199": "Hey, just curious, why 0.90+ QWK is unrealistic? ",
    "622211": "At least my CV scores were quite correlated with private and public. I observed good correlation when I started applying Ben's preprocessing:\n\nModel /  CV\t/   Public / Private\nBlend v3\t / 0.9328 / 0.813\t/ 0.924\nBlend v4\t /  0.9368 / 0.816 / 0.927\nBlend v5\t  /  0.9372 / 0.817 / 0.930\nBlend v6\t   /  0.9379 / 0.817 / 0.930",
    "621320": "@abhishek, I share most of your sentiments. You forgot to mention the unbelievably unfair submission issues.",
    "621196": "Agreed. The `train/test` distribution was a mess this time. ",
    "621334": "I jumped from 0.782 to 0.9... this is silly. Also I'm #1222 on the private LB but with a QWK difference wrt to the winner of barely 0.036... guess I'm satisfied with the result 👌 ",
    "621312": "Hey, Abhishek thanks for your kernels. I just started competing on kaggle and your work helped me a lot.  ",
    "621213": "why results are useless for the organizers ?",
    "621210": ""
  }
}