{
  "id": 22650,
  "title": "suggestions to make the next competition better",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22650",
  "author_name": "",
  "post_date": "2016-08-02T21:27:56.220Z",
  "votes": null,
  "comment_count": 5,
  "views": 484,
  "content": "<p>Thanks to all for making it a wonderful competition. I enjoy it a lot and learn a lot from it. It is time to do a quick review to see how we can improve it the next time. To start off with, here a my points:</p>\n\n<ol>\n<li><p>[test data]\ni though that a small amount of partial labelled test data wold be useful (e.g. 5 to 10%).This enables offline testing, etc</p></li>\n<li><p>[black box testing]\nthe final results can be based on totally  unseen data to prevent over fitting. or it can be graybox testing, e.g. 50% of the test data unseen</p></li>\n<li><p>[complexity]\nIf complexity is a target of the completion, flops can be used as part of the evaluation criteria. </p></li>\n<li><p>[use of test data] it would be better to state in advance if the test data can be used in training in any way, e.g. as soft targets in unsupervised ways</p></li>\n</ol>",
  "messages": [
    {
      "id": "129922",
      "postDate": "08/02/2016 21:27:56",
      "content": "<p>Thanks to all for making it a wonderful competition. I enjoy it a lot and learn a lot from it. It is time to do a quick review to see how we can improve it the next time. To start off with, here a my points:</p>\n\n<ol>\n<li><p>[test data]\ni though that a small amount of partial labelled test data wold be useful (e.g. 5 to 10%).This enables offline testing, etc</p></li>\n<li><p>[black box testing]\nthe final results can be based on totally  unseen data to prevent over fitting. or it can be graybox testing, e.g. 50% of the test data unseen</p></li>\n<li><p>[complexity]\nIf complexity is a target of the completion, flops can be used as part of the evaluation criteria. </p></li>\n<li><p>[use of test data] it would be better to state in advance if the test data can be used in training in any way, e.g. as soft targets in unsupervised ways</p></li>\n</ol>",
      "rawMarkdown": "Thanks to all for making it a wonderful competition. I enjoy it a lot and learn a lot from it. It is time to do a quick review to see how we can improve it the next time. To start off with, here a my points:\r\n\r\n 1. [test data]\r\ni though that a small amount of partial labelled test data wold be useful (e.g. 5 to 10%).This enables offline testing, etc\r\n\r\n 2. [black box testing]\r\nthe final results can be based on totally  unseen data to prevent over fitting. or it can be graybox testing, e.g. 50% of the test data unseen\r\n\r\n\r\n 3. [complexity]\r\nIf complexity is a target of the completion, flops can be used as part of the evaluation criteria. \r\n\r\n\r\n\r\n 4. [use of test data] it would be better to state in advance if the test data can be used in training in any way, e.g. as soft targets in unsupervised ways",
      "votes": null
    },
    {
      "id": "129937",
      "postDate": "08/02/2016 23:51:38",
      "content": "<p>Defiantly complexity  should be a criteria. Its useful for production systems and for these who have cheaper gpus. Hoverer its probably very hard to measure and verify.  </p>",
      "rawMarkdown": "Defiantly complexity  should be a criteria. Its useful for production systems and for these who have cheaper gpus. Hoverer its probably very hard to measure and verify.",
      "votes": null
    },
    {
      "id": "129939",
      "postDate": "08/03/2016 00:24:32",
      "content": "<p>Thank you for starting this post.</p>\n\n<p>I can't agree more with Heng Cherkeng's second suggestion. For this type of competition, private test data should not be given to the competitors at the beginning. They should not be available until, e.g., 24 hours (kyv's concern could also to some extent be solved) before the final deadline. Doing this could pretty much exclude the chance to perform any training to the real test data. This could also eliminate the cheaters who manually label and submit the test data. I always have the concern that someone could spend several days to manually label all the 79000+ test data, which is pretty light weight labor considering the time and efforts we put in this competition. They won't be caught unless they are greedy and reach top3.</p>",
      "rawMarkdown": "Thank you for starting this post.\r\n\r\nI can't agree more with Heng Cherkeng's second suggestion. For this type of competition, private test data should not be given to the competitors at the beginning. They should not be available until, e.g., 24 hours (kyv's concern could also to some extent be solved) before the final deadline. Doing this could pretty much exclude the chance to perform any training to the real test data. This could also eliminate the cheaters who manually label and submit the test data. I always have the concern that someone could spend several days to manually label all the 79000+ test data, which is pretty light weight labor considering the time and efforts we put in this competition. They won't be caught unless they are greedy and reach top3.",
      "votes": null
    },
    {
      "id": "129961",
      "postDate": "08/03/2016 03:44:57",
      "content": "<p>Guanshuo Xu's suggestion to not release private test data until 24 hours before the final deadline has some merit, but I think it would be unfair to those competitors who encounter some equipment failure or have to deal with some emergency on the last day of the contest and so would be disqualified despite having worked diligently up to that point.</p>",
      "rawMarkdown": "Guanshuo Xu's suggestion to not release private test data until 24 hours before the final deadline has some merit, but I think it would be unfair to those competitors who encounter some equipment failure or have to deal with some emergency on the last day of the contest and so would be disqualified despite having worked diligently up to that point.",
      "votes": null
    },
    {
      "id": "129968",
      "postDate": "08/03/2016 04:06:18",
      "content": "<p>@Guanshuo Xu, @David J. Slate\n.. Maybe you can try this solution:</p>\n\n<ul>\n<li>from day001 to day100: kaggle train their models. kagglers  test on some public dataset and submit results X</li>\n<li>from day101 to day115:models are to be fixed and cannot change. kagglers test on new released private dataset to submit results Y</li>\n</ul>\n\n<p>the competition results is based on Y. To verify if the model is not changed, submitted model should reproduce X.</p>",
      "rawMarkdown": "Guanshuo Xu, @David J. Slate\r\n.. Maybe you can try this solution:\r\n\r\n- from day001 to day100: kaggle train their models. kagglers  test on some public dataset and submit results X\r\n- from day101 to day115:models are to be fixed and cannot change. kagglers test on new released private dataset to submit results Y\r\n\r\nthe competition results is based on Y. To verify if the model is not changed, submitted model should reproduce X.",
      "votes": null
    },
    {
      "id": "129972",
      "postDate": "08/03/2016 05:07:55",
      "content": "<p>@Heng CherKeng, @Guanshuo Xu:</p>\n\n<p>I think the only real way to verify models would be to upload them so that Kaggle could run them on test data that the competitors have never seen.  Unfortunately models come in a variety of formats, are produced by a variety of machine learning applications and languages, and their performance may also depend on the operating system they are run on as well as the versions of libraries they depend on, and so on.  It would be difficult enough for Kaggle to try to do this just for the winners; doing it for all competitors would be a nightmarish task.</p>",
      "rawMarkdown": "Heng CherKeng, @Guanshuo Xu:\r\n\r\nI think the only real way to verify models would be to upload them so that Kaggle could run them on test data that the competitors have never seen.  Unfortunately models come in a variety of formats, are produced by a variety of machine learning applications and languages, and their performance may also depend on the operating system they are run on as well as the versions of libraries they depend on, and so on.  It would be difficult enough for Kaggle to try to do this just for the winners; doing it for all competitors would be a nightmarish task.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129937,
      "author_name": "yuraka",
      "author_url": "",
      "post_date": "08/02/2016 23:51:38",
      "content": "<p>Defiantly complexity  should be a criteria. Its useful for production systems and for these who have cheaper gpus. Hoverer its probably very hard to measure and verify.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129939,
      "author_name": "wowfattie",
      "author_url": "",
      "post_date": "08/03/2016 00:24:32",
      "content": "<p>Thank you for starting this post.</p>\n\n<p>I can't agree more with Heng Cherkeng's second suggestion. For this type of competition, private test data should not be given to the competitors at the beginning. They should not be available until, e.g., 24 hours (kyv's concern could also to some extent be solved) before the final deadline. Doing this could pretty much exclude the chance to perform any training to the real test data. This could also eliminate the cheaters who manually label and submit the test data. I always have the concern that someone could spend several days to manually label all the 79000+ test data, which is pretty light weight labor considering the time and efforts we put in this competition. They won't be caught unless they are greedy and reach top3.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129961,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "08/03/2016 03:44:57",
      "content": "<p>Guanshuo Xu's suggestion to not release private test data until 24 hours before the final deadline has some merit, but I think it would be unfair to those competitors who encounter some equipment failure or have to deal with some emergency on the last day of the contest and so would be disqualified despite having worked diligently up to that point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129968,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/03/2016 04:06:18",
      "content": "<p>@Guanshuo Xu, @David J. Slate\n.. Maybe you can try this solution:</p>\n\n<ul>\n<li>from day001 to day100: kaggle train their models. kagglers  test on some public dataset and submit results X</li>\n<li>from day101 to day115:models are to be fixed and cannot change. kagglers test on new released private dataset to submit results Y</li>\n</ul>\n\n<p>the competition results is based on Y. To verify if the model is not changed, submitted model should reproduce X.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129972,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "08/03/2016 05:07:55",
      "content": "<p>@Heng CherKeng, @Guanshuo Xu:</p>\n\n<p>I think the only real way to verify models would be to upload them so that Kaggle could run them on test data that the competitors have never seen.  Unfortunately models come in a variety of formats, are produced by a variety of machine learning applications and languages, and their performance may also depend on the operating system they are run on as well as the versions of libraries they depend on, and so on.  It would be difficult enough for Kaggle to try to do this just for the winners; doing it for all competitors would be a nightmarish task.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129922": "Thanks to all for making it a wonderful competition. I enjoy it a lot and learn a lot from it. It is time to do a quick review to see how we can improve it the next time. To start off with, here a my points:\r\n\r\n 1. [test data]\r\ni though that a small amount of partial labelled test data wold be useful (e.g. 5 to 10%).This enables offline testing, etc\r\n\r\n 2. [black box testing]\r\nthe final results can be based on totally  unseen data to prevent over fitting. or it can be graybox testing, e.g. 50% of the test data unseen\r\n\r\n\r\n 3. [complexity]\r\nIf complexity is a target of the completion, flops can be used as part of the evaluation criteria. \r\n\r\n\r\n\r\n 4. [use of test data] it would be better to state in advance if the test data can be used in training in any way, e.g. as soft targets in unsupervised ways",
    "129937": "Defiantly complexity  should be a criteria. Its useful for production systems and for these who have cheaper gpus. Hoverer its probably very hard to measure and verify.",
    "129939": "Thank you for starting this post.\r\n\r\nI can't agree more with Heng Cherkeng's second suggestion. For this type of competition, private test data should not be given to the competitors at the beginning. They should not be available until, e.g., 24 hours (kyv's concern could also to some extent be solved) before the final deadline. Doing this could pretty much exclude the chance to perform any training to the real test data. This could also eliminate the cheaters who manually label and submit the test data. I always have the concern that someone could spend several days to manually label all the 79000+ test data, which is pretty light weight labor considering the time and efforts we put in this competition. They won't be caught unless they are greedy and reach top3.",
    "129961": "Guanshuo Xu's suggestion to not release private test data until 24 hours before the final deadline has some merit, but I think it would be unfair to those competitors who encounter some equipment failure or have to deal with some emergency on the last day of the contest and so would be disqualified despite having worked diligently up to that point.",
    "129968": "Guanshuo Xu, @David J. Slate\r\n.. Maybe you can try this solution:\r\n\r\n- from day001 to day100: kaggle train their models. kagglers  test on some public dataset and submit results X\r\n- from day101 to day115:models are to be fixed and cannot change. kagglers test on new released private dataset to submit results Y\r\n\r\nthe competition results is based on Y. To verify if the model is not changed, submitted model should reproduce X.",
    "129972": "Heng CherKeng, @Guanshuo Xu:\r\n\r\nI think the only real way to verify models would be to upload them so that Kaggle could run them on test data that the competitors have never seen.  Unfortunately models come in a variety of formats, are produced by a variety of machine learning applications and languages, and their performance may also depend on the operating system they are run on as well as the versions of libraries they depend on, and so on.  It would be difficult enough for Kaggle to try to do this just for the winners; doing it for all competitors would be a nightmarish task."
  },
  "source": "meta"
}