{
  "id": 104992,
  "title": "We can train the model during submitting",
  "url": "/competitions/aptos2019-blindness-detection/discussion/104992",
  "author_name": "",
  "post_date": "2019-08-20T13:30:15.936656Z",
  "votes": null,
  "comment_count": 12,
  "views": 0,
  "content": "<p>While we submit your predicts, you can access the private test dataset. So I think that we can also tune our models by using the data of private test dataset.\nNeedless to say, we can not access the labels. But we can try the way like pseudo-labeling.\n<a href=\"https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\">https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969</a></p>\n\n<p>We have a GPU and 9 hour runtime. How you use it is up to you. \nWhat do you think about this？</p>",
  "messages": [
    {
      "id": "603596",
      "postDate": "08/20/2019 13:30:15",
      "content": "<p>While we submit your predicts, you can access the private test dataset. So I think that we can also tune our models by using the data of private test dataset.\nNeedless to say, we can not access the labels. But we can try the way like pseudo-labeling.\n<a href=\"https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\">https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969</a></p>\n\n<p>We have a GPU and 9 hour runtime. How you use it is up to you. \nWhat do you think about this？</p>",
      "rawMarkdown": "While we submit your predicts, you can access the private test dataset. So I think that we can also tune our models by using the data of private test dataset.\nNeedless to say, we can not access the labels. But we can try the way like pseudo-labeling.\nhttps://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\n\nWe have a GPU and 9 hour runtime. How you use it is up to you. \nWhat do you think about this？",
      "votes": null
    },
    {
      "id": "603651",
      "postDate": "08/20/2019 14:22:14",
      "content": "<p>There is a reason why this competition is kernel-only :)). Can this method apply when deploying for the product?.  Real victory only if your solution can be used by the organizer in reality, I think.</p>",
      "rawMarkdown": "There is a reason why this competition is kernel-only :)). Can this method apply when deploying for the product?.  Real victory only if your solution can be used by the organizer in reality, I think.",
      "votes": null
    },
    {
      "id": "603687",
      "postDate": "08/20/2019 14:56:48",
      "content": "<p>Now I can ensemble 150 models and it takes 7 hours for submit. Such impractical solutions have won in many competitions so far. If the kaggle team doesn't have rules that prohibit it, I think it's allowed. </p>",
      "rawMarkdown": "Now I can ensemble 150 models and it takes 7 hours for submit. Such impractical solutions have won in many competitions so far. If the kaggle team doesn't have rules that prohibit it, I think it's allowed.",
      "votes": null
    },
    {
      "id": "603694",
      "postDate": "08/20/2019 15:00:27",
      "content": "<p>yes, its allowed :), i only say a few personal views ^^.</p>",
      "rawMarkdown": "yes, its allowed :), i only say a few personal views ^^.",
      "votes": null
    },
    {
      "id": "603697",
      "postDate": "08/20/2019 15:07:45",
      "content": "<p>Can you show me the trick of how to ensemble 150 models with this amount of data?</p>",
      "rawMarkdown": "Can you show me the trick of how to ensemble 150 models with this amount of data?",
      "votes": null
    },
    {
      "id": "603701",
      "postDate": "08/20/2019 15:11:20",
      "content": "<p>I also think a simple and reasonable solution is cool.\nBy the way, what do you think about training while submitting by using pseudo-labeling?</p>",
      "rawMarkdown": "I also think a simple and reasonable solution is cool.\nBy the way, what do you think about training while submitting by using pseudo-labeling?",
      "votes": null
    },
    {
      "id": "603702",
      "postDate": "08/20/2019 15:14:07",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I mean inferring 150 times for each image. The trick is to do all the data pre-processing first. But it did not improve my score ... Using too many models is not good solution.</p>",
      "rawMarkdown": "philippsinger I mean inferring 150 times for each image. The trick is to do all the data pre-processing first. But it did not improve my score ... Using too many models is not good solution.",
      "votes": null
    },
    {
      "id": "603709",
      "postDate": "08/20/2019 15:23:05",
      "content": "<p>I'm not sure between the trade-off time to train pseudo labeling and using more models to ensemble :). Now, my single model obtained 0.847 in public LB (no tta, no optimize threshold).</p>",
      "rawMarkdown": "I'm not sure between the trade-off time to train pseudo labeling and using more models to ensemble :). Now, my single model obtained 0.847 in public LB (no tta, no optimize threshold).",
      "votes": null
    },
    {
      "id": "603724",
      "postDate": "08/20/2019 15:43:41",
      "content": "<p>I doubt that you can do it for full test data though <a href=\"/hirune924\">@hirune924</a> within 9 hours 150 times</p>",
      "rawMarkdown": "I doubt that you can do it for full test data though @hirune924 within 9 hours 150 times",
      "votes": null
    },
    {
      "id": "603738",
      "postDate": "08/20/2019 15:55:31",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Do you use n-fold and infer n times? This difference in feeling may come from the difference in how the models are counted</p>",
      "rawMarkdown": "philippsinger Do you use n-fold and infer n times? This difference in feeling may come from the difference in how the models are counted",
      "votes": null
    },
    {
      "id": "603753",
      "postDate": "08/20/2019 16:11:41",
      "content": "<p>I am just saying that doing inference 150 times on ~13k images will be tough. </p>",
      "rawMarkdown": "I am just saying that doing inference 150 times on ~13k images will be tough.",
      "votes": null
    },
    {
      "id": "603768",
      "postDate": "08/20/2019 16:27:25",
      "content": "<p>do you use such heavy models? tell me your inference performance.</p>",
      "rawMarkdown": "do you use such heavy models? tell me your inference performance.",
      "votes": null
    },
    {
      "id": "604301",
      "postDate": "08/21/2019 08:45:32",
      "content": "<p>Don't know about you but tuning model actually reduces model accuracy (At least in my case).</p>",
      "rawMarkdown": "Don't know about you but tuning model actually reduces model accuracy (At least in my case).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 603651,
      "author_name": "dathudeptrai",
      "author_url": "",
      "post_date": "08/20/2019 14:22:14",
      "content": "<p>There is a reason why this competition is kernel-only :)). Can this method apply when deploying for the product?.  Real victory only if your solution can be used by the organizer in reality, I think.</p>",
      "votes": null,
      "replies": [
        {
          "id": 603687,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/20/2019 14:56:48",
          "content": "<p>Now I can ensemble 150 models and it takes 7 hours for submit. Such impractical solutions have won in many competitions so far. If the kaggle team doesn't have rules that prohibit it, I think it's allowed. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603694,
          "author_name": "dathudeptrai",
          "author_url": "",
          "post_date": "08/20/2019 15:00:27",
          "content": "<p>yes, its allowed :), i only say a few personal views ^^.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603697,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/20/2019 15:07:45",
          "content": "<p>Can you show me the trick of how to ensemble 150 models with this amount of data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603701,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/20/2019 15:11:20",
          "content": "<p>I also think a simple and reasonable solution is cool.\nBy the way, what do you think about training while submitting by using pseudo-labeling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603702,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/20/2019 15:14:07",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I mean inferring 150 times for each image. The trick is to do all the data pre-processing first. But it did not improve my score ... Using too many models is not good solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603709,
          "author_name": "dathudeptrai",
          "author_url": "",
          "post_date": "08/20/2019 15:23:05",
          "content": "<p>I'm not sure between the trade-off time to train pseudo labeling and using more models to ensemble :). Now, my single model obtained 0.847 in public LB (no tta, no optimize threshold).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603724,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/20/2019 15:43:41",
          "content": "<p>I doubt that you can do it for full test data though <a href=\"/hirune924\">@hirune924</a> within 9 hours 150 times</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603738,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/20/2019 15:55:31",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Do you use n-fold and infer n times? This difference in feeling may come from the difference in how the models are counted</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603753,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/20/2019 16:11:41",
          "content": "<p>I am just saying that doing inference 150 times on ~13k images will be tough. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 603768,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/20/2019 16:27:25",
          "content": "<p>do you use such heavy models? tell me your inference performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 604301,
      "author_name": "quandapro",
      "author_url": "",
      "post_date": "08/21/2019 08:45:32",
      "content": "<p>Don't know about you but tuning model actually reduces model accuracy (At least in my case).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "603596": "While we submit your predicts, you can access the private test dataset. So I think that we can also tune our models by using the data of private test dataset.\nNeedless to say, we can not access the labels. But we can try the way like pseudo-labeling.\nhttps://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\n\nWe have a GPU and 9 hour runtime. How you use it is up to you. \nWhat do you think about this？",
    "603651": "There is a reason why this competition is kernel-only :)). Can this method apply when deploying for the product?.  Real victory only if your solution can be used by the organizer in reality, I think.",
    "603687": "Now I can ensemble 150 models and it takes 7 hours for submit. Such impractical solutions have won in many competitions so far. If the kaggle team doesn't have rules that prohibit it, I think it's allowed.",
    "603694": "yes, its allowed :), i only say a few personal views ^^.",
    "603697": "Can you show me the trick of how to ensemble 150 models with this amount of data?",
    "603701": "I also think a simple and reasonable solution is cool.\nBy the way, what do you think about training while submitting by using pseudo-labeling?",
    "603702": "philippsinger I mean inferring 150 times for each image. The trick is to do all the data pre-processing first. But it did not improve my score ... Using too many models is not good solution.",
    "603709": "I'm not sure between the trade-off time to train pseudo labeling and using more models to ensemble :). Now, my single model obtained 0.847 in public LB (no tta, no optimize threshold).",
    "603724": "I doubt that you can do it for full test data though @hirune924 within 9 hours 150 times",
    "603738": "philippsinger Do you use n-fold and infer n times? This difference in feeling may come from the difference in how the models are counted",
    "603753": "I am just saying that doing inference 150 times on ~13k images will be tough.",
    "603768": "do you use such heavy models? tell me your inference performance.",
    "604301": "Don't know about you but tuning model actually reduces model accuracy (At least in my case)."
  },
  "source": "meta"
}