{
  "id": 108089,
  "title": "Open Letter",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108089",
  "author_name": "",
  "post_date": "2019-09-09T02:13:25.896648900Z",
  "votes": -15,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Dear the competition organizers,</p>\n\n<p>Thanks for your hard work in organizing the competition APTOS 2019 Blindness Detection！ </p>\n\n<p>But we were ranked 192, which is obviously very disappointing because we were ranked 52 before the final submission.  According to the competition rules, two models that performed best on public test data were submitted and they were run on private test data. Then, our ranking went down from 52 to 192 - down by 140!</p>\n\n<p>We do not have the right to challenge the rules. However, such rules encourage random ranking instead of hard work. This can be seen in other teams, for example, the team ranked 24 went straight down to 291, and the team ranked 2nd went down to 83, even some teams went down more than 1000 !</p>\n\n<p>What's more, Kaggle has already run all submitted models on private data, and our best model on private data can achieve 0.922 (currently 0.918). Then, why didnot Kaggle use the best model?</p>\n\n<p>There are many unreasonable places in this game. Hope Kaggle and the organizers can review the rules, and we look forward to taking more competitions in future. </p>",
  "messages": [
    {
      "id": "621835",
      "postDate": "09/09/2019 02:13:25",
      "content": "<p>Dear the competition organizers,</p>\n\n<p>Thanks for your hard work in organizing the competition APTOS 2019 Blindness Detection！ </p>\n\n<p>But we were ranked 192, which is obviously very disappointing because we were ranked 52 before the final submission.  According to the competition rules, two models that performed best on public test data were submitted and they were run on private test data. Then, our ranking went down from 52 to 192 - down by 140!</p>\n\n<p>We do not have the right to challenge the rules. However, such rules encourage random ranking instead of hard work. This can be seen in other teams, for example, the team ranked 24 went straight down to 291, and the team ranked 2nd went down to 83, even some teams went down more than 1000 !</p>\n\n<p>What's more, Kaggle has already run all submitted models on private data, and our best model on private data can achieve 0.922 (currently 0.918). Then, why didnot Kaggle use the best model?</p>\n\n<p>There are many unreasonable places in this game. Hope Kaggle and the organizers can review the rules, and we look forward to taking more competitions in future. </p>",
      "rawMarkdown": "Dear the competition organizers,\n\nThanks for your hard work in organizing the competition APTOS 2019 Blindness Detection！ \n\nBut we were ranked 192, which is obviously very disappointing because we were ranked 52 before the final submission.  According to the competition rules, two models that performed best on public test data were submitted and they were run on private test data. Then, our ranking went down from 52 to 192 - down by 140!\n\nWe do not have the right to challenge the rules. However, such rules encourage random ranking instead of hard work. This can be seen in other teams, for example, the team ranked 24 went straight down to 291, and the team ranked 2nd went down to 83, even some teams went down more than 1000 !\n\nWhat's more, Kaggle has already run all submitted models on private data, and our best model on private data can achieve 0.922 (currently 0.918). Then, why didnot Kaggle use the best model?\n\nThere are many unreasonable places in this game. Hope Kaggle and the organizers can review the rules, and we look forward to taking more competitions in future.",
      "votes": null
    },
    {
      "id": "621840",
      "postDate": "09/09/2019 02:23:52",
      "content": "<p>Since we are mentioned (3 &gt;&gt; 83) , I will say that in our case, we were unable to judge best robust submission. So it’s our own fault, and we are looking forward to improve next time.</p>",
      "rawMarkdown": "Since we are mentioned (3 &gt;&gt; 83) , I will say that in our case, we were unable to judge best robust submission. So it’s our own fault, and we are looking forward to improve next time.",
      "votes": null
    },
    {
      "id": "621851",
      "postDate": "09/09/2019 02:57:48",
      "content": "<p>Yea, something lucky.</p>",
      "rawMarkdown": "Yea, something lucky.",
      "votes": null
    },
    {
      "id": "621893",
      "postDate": "09/09/2019 04:38:56",
      "content": "<p>If Kaggle can auto choose the best private score for us as final score, I'll just do 30+ times ensemble with different models, and use more tricks to overfit public LB without concerning anything ;)</p>",
      "rawMarkdown": "If Kaggle can auto choose the best private score for us as final score, I'll just do 30+ times ensemble with different models, and use more tricks to overfit public LB without concerning anything ;)",
      "votes": null
    },
    {
      "id": "621909",
      "postDate": "09/09/2019 04:53:14",
      "content": "<p>You still have to be concerned about one problem =))), the inference time :)). The contest will be more interesting if the execution time of each submission is limited to only one or two models. Ensemble many models also can not be applied in practice. BTW, my best single model is 92.8 on private lb but I chose the model with only 90.0 instead of 92.8 :( very sad but this is my fault: D</p>",
      "rawMarkdown": "You still have to be concerned about one problem =))), the inference time :)). The contest will be more interesting if the execution time of each submission is limited to only one or two models. Ensemble many models also can not be applied in practice. BTW, my best single model is 92.8 on private lb but I chose the model with only 90.0 instead of 92.8 :( very sad but this is my fault: D",
      "votes": null
    },
    {
      "id": "621913",
      "postDate": "09/09/2019 04:58:03",
      "content": "<p>I mean, say, I have 15 models and I will randomly pick up 6~7 of them to do ensemble, repeat 30+ times. Then just sit there, grab a coffee and wait for Kaggle to find out which is the <strong>best ensemble</strong> for the private set.\nAnd, sorry to hear that...</p>",
      "rawMarkdown": "I mean, say, I have 15 models and I will randomly pick up 6~7 of them to do ensemble, repeat 30+ times. Then just sit there, grab a coffee and wait for Kaggle to find out which is the **best ensemble** for the private set.\nAnd, sorry to hear that...",
      "votes": null
    },
    {
      "id": "621940",
      "postDate": "09/09/2019 05:34:35",
      "content": "<p>In real world scenarios we never get to see test data (Images + labels) while selecting our best models before putting it to production. We have to select our best model based on val set performance or hold-out test set performances. So, Kaggle's rule of selecting your 2 most faithful models as final models replicates those real world scenarios. It makes sense.</p>",
      "rawMarkdown": "In real world scenarios we never get to see test data (Images + labels) while selecting our best models before putting it to production. We have to select our best model based on val set performance or hold-out test set performances. So, Kaggle's rule of selecting your 2 most faithful models as final models replicates those real world scenarios. It makes sense.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 621840,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "09/09/2019 02:23:52",
      "content": "<p>Since we are mentioned (3 &gt;&gt; 83) , I will say that in our case, we were unable to judge best robust submission. So it’s our own fault, and we are looking forward to improve next time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 621851,
          "author_name": "lifengnan",
          "author_url": "",
          "post_date": "09/09/2019 02:57:48",
          "content": "<p>Yea, something lucky.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 621893,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "09/09/2019 04:38:56",
      "content": "<p>If Kaggle can auto choose the best private score for us as final score, I'll just do 30+ times ensemble with different models, and use more tricks to overfit public LB without concerning anything ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 621909,
          "author_name": "dathudeptrai",
          "author_url": "",
          "post_date": "09/09/2019 04:53:14",
          "content": "<p>You still have to be concerned about one problem =))), the inference time :)). The contest will be more interesting if the execution time of each submission is limited to only one or two models. Ensemble many models also can not be applied in practice. BTW, my best single model is 92.8 on private lb but I chose the model with only 90.0 instead of 92.8 :( very sad but this is my fault: D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 621913,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/09/2019 04:58:03",
          "content": "<p>I mean, say, I have 15 models and I will randomly pick up 6~7 of them to do ensemble, repeat 30+ times. Then just sit there, grab a coffee and wait for Kaggle to find out which is the <strong>best ensemble</strong> for the private set.\nAnd, sorry to hear that...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 621940,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "09/09/2019 05:34:35",
      "content": "<p>In real world scenarios we never get to see test data (Images + labels) while selecting our best models before putting it to production. We have to select our best model based on val set performance or hold-out test set performances. So, Kaggle's rule of selecting your 2 most faithful models as final models replicates those real world scenarios. It makes sense.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621835": "Dear the competition organizers,\n\nThanks for your hard work in organizing the competition APTOS 2019 Blindness Detection！ \n\nBut we were ranked 192, which is obviously very disappointing because we were ranked 52 before the final submission.  According to the competition rules, two models that performed best on public test data were submitted and they were run on private test data. Then, our ranking went down from 52 to 192 - down by 140!\n\nWe do not have the right to challenge the rules. However, such rules encourage random ranking instead of hard work. This can be seen in other teams, for example, the team ranked 24 went straight down to 291, and the team ranked 2nd went down to 83, even some teams went down more than 1000 !\n\nWhat's more, Kaggle has already run all submitted models on private data, and our best model on private data can achieve 0.922 (currently 0.918). Then, why didnot Kaggle use the best model?\n\nThere are many unreasonable places in this game. Hope Kaggle and the organizers can review the rules, and we look forward to taking more competitions in future.",
    "621840": "Since we are mentioned (3 &gt;&gt; 83) , I will say that in our case, we were unable to judge best robust submission. So it’s our own fault, and we are looking forward to improve next time.",
    "621851": "Yea, something lucky.",
    "621893": "If Kaggle can auto choose the best private score for us as final score, I'll just do 30+ times ensemble with different models, and use more tricks to overfit public LB without concerning anything ;)",
    "621909": "You still have to be concerned about one problem =))), the inference time :)). The contest will be more interesting if the execution time of each submission is limited to only one or two models. Ensemble many models also can not be applied in practice. BTW, my best single model is 92.8 on private lb but I chose the model with only 90.0 instead of 92.8 :( very sad but this is my fault: D",
    "621913": "I mean, say, I have 15 models and I will randomly pick up 6~7 of them to do ensemble, repeat 30+ times. Then just sit there, grab a coffee and wait for Kaggle to find out which is the **best ensemble** for the private set.\nAnd, sorry to hear that...",
    "621940": "In real world scenarios we never get to see test data (Images + labels) while selecting our best models before putting it to production. We have to select our best model based on val set performance or hold-out test set performances. So, Kaggle's rule of selecting your 2 most faithful models as final models replicates those real world scenarios. It makes sense."
  },
  "source": "meta"
}