{
  "id": 43667,
  "title": "Whats stopping a team from manually annotating the test set?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43667",
  "author_name": "",
  "post_date": "2017-11-17T15:55:43.818051800Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Whats stopping a team from manually annotating the test set?</p>\n\n<p>I know its a lot of files (&gt;150k). But in principle a group with <em>alot</em> of patience could annotate those files for which their classifier output a low confidence on their estimates (see for instance <a href=\"http://mlg.eng.cam.ac.uk/yarin/blog_3d801aa532c1ce.html\">http://mlg.eng.cam.ac.uk/yarin/blog_3d801aa532c1ce.html</a> ). </p>\n\n<p>I think the best approach would be to require people to submit their code and run it on a hidden test set in a sand-boxed environment. </p>",
  "messages": [
    {
      "id": "245077",
      "postDate": "11/17/2017 15:55:43",
      "content": "<p>Whats stopping a team from manually annotating the test set?</p>\n\n<p>I know its a lot of files (&gt;150k). But in principle a group with <em>alot</em> of patience could annotate those files for which their classifier output a low confidence on their estimates (see for instance <a href=\"http://mlg.eng.cam.ac.uk/yarin/blog_3d801aa532c1ce.html\">http://mlg.eng.cam.ac.uk/yarin/blog_3d801aa532c1ce.html</a> ). </p>\n\n<p>I think the best approach would be to require people to submit their code and run it on a hidden test set in a sand-boxed environment. </p>",
      "rawMarkdown": "Whats stopping a team from manually annotating the test set?\n\nI know its a lot of files (&gt;150k). But in principle a group with *alot* of patience could annotate those files for which their classifier output a low confidence on their estimates (see for instance http://mlg.eng.cam.ac.uk/yarin/blog_3d801aa532c1ce.html ). \n\n\nI think the best approach would be to require people to submit their code and run it on a hidden test set in a sand-boxed environment.",
      "votes": null
    },
    {
      "id": "245081",
      "postDate": "11/17/2017 16:01:22",
      "content": "<p>Hi Ptolemy,</p>\n\n<p>Just a note that listed in the rules you agreed to is the following. I hope that helps!</p>\n\n<p>\"As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:</p>\n\n<p>(a) deliver to the Competition Sponsor the final model’s software code as used to generate the winning Submission and associated documentation. The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;\"</p>",
      "rawMarkdown": "Hi Ptolemy,\n\nJust a note that listed in the rules you agreed to is the following. I hope that helps!\n\n\"As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:\n\n(a) deliver to the Competition Sponsor the final model’s software code as used to generate the winning Submission and associated documentation. The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;\"",
      "votes": null
    },
    {
      "id": "245100",
      "postDate": "11/17/2017 16:35:52",
      "content": "<p>They only need to deliver the trained Neural Network. There's no way for the judges to know that the trained neural network used labeled test data. </p>",
      "rawMarkdown": "They only need to deliver the trained Neural Network. There's no way for the judges to know that the trained neural network used labeled test data.",
      "votes": null
    },
    {
      "id": "245171",
      "postDate": "11/17/2017 18:42:41",
      "content": "<p>No the rules say model's software the code that generates the solution. That is not the same as the \"trained Neural Network\". If the code was stupid enough to use labeled test data, its invalid. How are you going to hide this fact in code? The whole thing must be repeatable. And if you after training delete the part where it was trained over test labels, the results won't be reproducible.</p>",
      "rawMarkdown": "No the rules say model's software the code that generates the solution. That is not the same as the \"trained Neural Network\". If the code was stupid enough to use labeled test data, its invalid. How are you going to hide this fact in code? The whole thing must be repeatable. And if you after training delete the part where it was trained over test labels, the results won't be reproducible.",
      "votes": null
    },
    {
      "id": "245194",
      "postDate": "11/17/2017 19:16:26",
      "content": "<p>\"model’s software code as used to generate the winning Submission\"</p>\n\n<p>Sounds the same as trained model to me. \nOtherwise, good luck trying to reproduce results from classifiers that rely on stochastic seeding and training (including Neural Nets)! That would be even easier for the cheater to claim that differences in performance is due to stochastic nature of the initialization and training and he was able to \"stumble\" on a good one during the competition. </p>\n\n<p>The point is that he does not have to label all the data, only the hard ones (ie, the support vector) from the test set. And this can give him/her a significant advantage over others in developing his classifier and understanding the problem. </p>",
      "rawMarkdown": "\"model’s software code as used to generate the winning Submission\"\n\nSounds the same as trained model to me. \nOtherwise, good luck trying to reproduce results from classifiers that rely on stochastic seeding and training (including Neural Nets)! That would be even easier for the cheater to claim that differences in performance is due to stochastic nature of the initialization and training and he was able to \"stumble\" on a good one during the competition. \n\nThe point is that he does not have to label all the data, only the hard ones (ie, the support vector) from the test set. And this can give him/her a significant advantage over others in developing his classifier and understanding the problem.",
      "votes": null
    },
    {
      "id": "245198",
      "postDate": "11/17/2017 19:29:34",
      "content": "<p>people have submitted human determined labels to ILSVRC in the past :-)</p>\n\n<p><a href=\"http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/\">http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/</a></p>",
      "rawMarkdown": "people have submitted human determined labels to ILSVRC in the past :-)\n\nhttp://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 245081,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "11/17/2017 16:01:22",
      "content": "<p>Hi Ptolemy,</p>\n\n<p>Just a note that listed in the rules you agreed to is the following. I hope that helps!</p>\n\n<p>\"As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:</p>\n\n<p>(a) deliver to the Competition Sponsor the final model’s software code as used to generate the winning Submission and associated documentation. The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 245100,
          "author_name": "ptolemy",
          "author_url": "",
          "post_date": "11/17/2017 16:35:52",
          "content": "<p>They only need to deliver the trained Neural Network. There's no way for the judges to know that the trained neural network used labeled test data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245171,
          "author_name": "shiqbal",
          "author_url": "",
          "post_date": "11/17/2017 18:42:41",
          "content": "<p>No the rules say model's software the code that generates the solution. That is not the same as the \"trained Neural Network\". If the code was stupid enough to use labeled test data, its invalid. How are you going to hide this fact in code? The whole thing must be repeatable. And if you after training delete the part where it was trained over test labels, the results won't be reproducible.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245194,
          "author_name": "ptolemy",
          "author_url": "",
          "post_date": "11/17/2017 19:16:26",
          "content": "<p>\"model’s software code as used to generate the winning Submission\"</p>\n\n<p>Sounds the same as trained model to me. \nOtherwise, good luck trying to reproduce results from classifiers that rely on stochastic seeding and training (including Neural Nets)! That would be even easier for the cheater to claim that differences in performance is due to stochastic nature of the initialization and training and he was able to \"stumble\" on a good one during the competition. </p>\n\n<p>The point is that he does not have to label all the data, only the hard ones (ie, the support vector) from the test set. And this can give him/her a significant advantage over others in developing his classifier and understanding the problem. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 245198,
      "author_name": "holzner",
      "author_url": "",
      "post_date": "11/17/2017 19:29:34",
      "content": "<p>people have submitted human determined labels to ILSVRC in the past :-)</p>\n\n<p><a href=\"http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/\">http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "245077": "Whats stopping a team from manually annotating the test set?\n\nI know its a lot of files (&gt;150k). But in principle a group with *alot* of patience could annotate those files for which their classifier output a low confidence on their estimates (see for instance http://mlg.eng.cam.ac.uk/yarin/blog_3d801aa532c1ce.html ). \n\n\nI think the best approach would be to require people to submit their code and run it on a hidden test set in a sand-boxed environment.",
    "245081": "Hi Ptolemy,\n\nJust a note that listed in the rules you agreed to is the following. I hope that helps!\n\n\"As a condition to being awarded a Prize, a Prize winner must fulfill the following obligations:\n\n(a) deliver to the Competition Sponsor the final model’s software code as used to generate the winning Submission and associated documentation. The delivered software code must be capable of generating the winning Submission and contain a description of resources required to build and/or run the executable code successfully;\"",
    "245100": "They only need to deliver the trained Neural Network. There's no way for the judges to know that the trained neural network used labeled test data.",
    "245171": "No the rules say model's software the code that generates the solution. That is not the same as the \"trained Neural Network\". If the code was stupid enough to use labeled test data, its invalid. How are you going to hide this fact in code? The whole thing must be repeatable. And if you after training delete the part where it was trained over test labels, the results won't be reproducible.",
    "245194": "\"model’s software code as used to generate the winning Submission\"\n\nSounds the same as trained model to me. \nOtherwise, good luck trying to reproduce results from classifiers that rely on stochastic seeding and training (including Neural Nets)! That would be even easier for the cheater to claim that differences in performance is due to stochastic nature of the initialization and training and he was able to \"stumble\" on a good one during the competition. \n\nThe point is that he does not have to label all the data, only the hard ones (ie, the support vector) from the test set. And this can give him/her a significant advantage over others in developing his classifier and understanding the problem.",
    "245198": "people have submitted human determined labels to ILSVRC in the past :-)\n\nhttp://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/"
  },
  "source": "meta"
}