{
  "id": 20335,
  "title": "Crowdsource annotation of training set",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20335",
  "author_name": "Toby Cheese",
  "post_date": "2016-04-22T13:37:24.587000",
  "votes": 5,
  "comment_count": 5,
  "views": 1473,
  "content": "<p>Hi all</p>\n\n<p>I hacked together a web application so we can work on annotating the training set. You find the first version <a href=\"http://tobycheese.pythonanywhere.com/\">here</a>. Before launching the effort for real, I need some feedback on a few things though:</p>\n\n<p><strong>What parts of the image should we annotate</strong>?\nAs a start, I configured three ROIs per image: head, steering wheel and left arm. But we need a consensus as which parts we want to annotate before starting the effort.</p>\n\n<p><strong>In what order should the images appear for tagging?</strong>\nAt the moment, the images are given in the order they are globbed from the filesystem. But this is suboptimal. One suggestion would be to bundle images per driver and in consequence to allow only as many people as there are distinct drivers to work in parallel. When doing it this way, for each new image the previously made annotations could be suggested, which would speed up annotations for a sequence of very similar images considerably. </p>\n\n<p><strong>How could handling be improved?</strong>\nPlease play with it a little, and report ideas on how the annotation process can be more speedy or more convenient.</p>\n\n<p><strong>Bugs?</strong>\nThey are somewhere, please report any sightings.</p>\n\n<p><strong>How should the (final) output be formatted?</strong>\nThis is maybe less important, but at the moment output is simply all fields of the database dumped into json. If someone has a good suggestion, I can make output more convenient.</p>\n\n<p><strong>Some more notes:</strong></p>\n\n<p>To access the results, use the following two links to either get the results of one single image, or of all images annotated so far:\n<a href=\"http://tobycheese.pythonanywhere.com/annotated/img_72999\">http://tobycheese.pythonanywhere.com/annotated/img_72999</a>\n<a href=\"http://tobycheese.pythonanywhere.com/annotated/\">http://tobycheese.pythonanywhere.com/annotated</a></p>\n\n<p>If you have already spent a substantial amount of work or money on annotating images and want to contribute this, we will find a way. At first I need all image ids that you have already annotated, so I can filter them out to avoid double work. Then (after we have decided on which ROIs to annotate), I can set up a second instance of the web application with just your images to annotate the missing ROIS (in case the ROIs you annotated differ from those here). Then your work can be integrated here.</p>\n\n<p><strong>Things still on the todo list:</strong></p>\n\n<ul>\n<li>Get answers to above questions and implement the consequences</li>\n<li>Add keyboard shortcuts for 'Next' and 'Reset Annotations'</li>\n<li>Add a cleanup process to remove locked files</li>\n<li>Make it SQL injection safe</li>\n<li>Maybe make it look a bit nicer, it really is just a quick hack</li>\n<li>Polish code and put it on github</li>\n</ul>\n\n<hr>\n\n<p>Edit:</p>\n\n<p><strong>Changes since initial post:</strong></p>\n\n<ul>\n<li>Reset button actually does something now</li>\n<li>Added a reference to libraries used</li>\n</ul>",
  "messages": [
    {
      "id": 117150,
      "postDate": "2016-04-27T15:46:27.187Z",
      "content": "<p>@Toby -</p>\n\n<p>I'm taking a guess here. Either (a) people think hand annotations will be too much work, and might not give incremental benefit over, e.g., pre-trained networks, or (b) people are annotating their images themselves to keep it as a competitive advantage.</p>",
      "rawMarkdown": "@Toby -\r\n\r\nI'm taking a guess here. Either (a) people think hand annotations will be too much work, and might not give incremental benefit over, e.g., pre-trained networks, or (b) people are annotating their images themselves to keep it as a competitive advantage.",
      "votes": 3
    },
    {
      "id": 116193,
      "postDate": "2016-04-22T13:37:24.587Z",
      "content": "<p>Hi all</p>\n\n<p>I hacked together a web application so we can work on annotating the training set. You find the first version <a href=\"http://tobycheese.pythonanywhere.com/\">here</a>. Before launching the effort for real, I need some feedback on a few things though:</p>\n\n<p><strong>What parts of the image should we annotate</strong>?\nAs a start, I configured three ROIs per image: head, steering wheel and left arm. But we need a consensus as which parts we want to annotate before starting the effort.</p>\n\n<p><strong>In what order should the images appear for tagging?</strong>\nAt the moment, the images are given in the order they are globbed from the filesystem. But this is suboptimal. One suggestion would be to bundle images per driver and in consequence to allow only as many people as there are distinct drivers to work in parallel. When doing it this way, for each new image the previously made annotations could be suggested, which would speed up annotations for a sequence of very similar images considerably. </p>\n\n<p><strong>How could handling be improved?</strong>\nPlease play with it a little, and report ideas on how the annotation process can be more speedy or more convenient.</p>\n\n<p><strong>Bugs?</strong>\nThey are somewhere, please report any sightings.</p>\n\n<p><strong>How should the (final) output be formatted?</strong>\nThis is maybe less important, but at the moment output is simply all fields of the database dumped into json. If someone has a good suggestion, I can make output more convenient.</p>\n\n<p><strong>Some more notes:</strong></p>\n\n<p>To access the results, use the following two links to either get the results of one single image, or of all images annotated so far:\n<a href=\"http://tobycheese.pythonanywhere.com/annotated/img_72999\">http://tobycheese.pythonanywhere.com/annotated/img_72999</a>\n<a href=\"http://tobycheese.pythonanywhere.com/annotated/\">http://tobycheese.pythonanywhere.com/annotated</a></p>\n\n<p>If you have already spent a substantial amount of work or money on annotating images and want to contribute this, we will find a way. At first I need all image ids that you have already annotated, so I can filter them out to avoid double work. Then (after we have decided on which ROIs to annotate), I can set up a second instance of the web application with just your images to annotate the missing ROIS (in case the ROIs you annotated differ from those here). Then your work can be integrated here.</p>\n\n<p><strong>Things still on the todo list:</strong></p>\n\n<ul>\n<li>Get answers to above questions and implement the consequences</li>\n<li>Add keyboard shortcuts for 'Next' and 'Reset Annotations'</li>\n<li>Add a cleanup process to remove locked files</li>\n<li>Make it SQL injection safe</li>\n<li>Maybe make it look a bit nicer, it really is just a quick hack</li>\n<li>Polish code and put it on github</li>\n</ul>\n\n<hr>\n\n<p>Edit:</p>\n\n<p><strong>Changes since initial post:</strong></p>\n\n<ul>\n<li>Reset button actually does something now</li>\n<li>Added a reference to libraries used</li>\n</ul>",
      "rawMarkdown": "Hi all\r\n\r\nI hacked together a web application so we can work on annotating the training set. You find the first version [here][1]. Before launching the effort for real, I need some feedback on a few things though:\r\n\r\n\r\n**What parts of the image should we annotate**?\r\nAs a start, I configured three ROIs per image: head, steering wheel and left arm. But we need a consensus as which parts we want to annotate before starting the effort.\r\n\r\n\r\n**In what order should the images appear for tagging?**\r\nAt the moment, the images are given in the order they are globbed from the filesystem. But this is suboptimal. One suggestion would be to bundle images per driver and in consequence to allow only as many people as there are distinct drivers to work in parallel. When doing it this way, for each new image the previously made annotations could be suggested, which would speed up annotations for a sequence of very similar images considerably. \r\n\r\n**How could handling be improved?**\r\nPlease play with it a little, and report ideas on how the annotation process can be more speedy or more convenient.\r\n\r\n**Bugs?**\r\nThey are somewhere, please report any sightings.\r\n\r\n**How should the (final) output be formatted?**\r\nThis is maybe less important, but at the moment output is simply all fields of the database dumped into json. If someone has a good suggestion, I can make output more convenient.\r\n\r\n\r\n**Some more notes:**\r\n\r\nTo access the results, use the following two links to either get the results of one single image, or of all images annotated so far:\r\n[http://tobycheese.pythonanywhere.com/annotated/img_72999][2]\r\n[http://tobycheese.pythonanywhere.com/annotated][3]\r\n\r\nIf you have already spent a substantial amount of work or money on annotating images and want to contribute this, we will find a way. At first I need all image ids that you have already annotated, so I can filter them out to avoid double work. Then (after we have decided on which ROIs to annotate), I can set up a second instance of the web application with just your images to annotate the missing ROIS (in case the ROIs you annotated differ from those here). Then your work can be integrated here.\r\n\r\n**Things still on the todo list:**\r\n\r\n - Get answers to above questions and implement the consequences\r\n - Add keyboard shortcuts for 'Next' and 'Reset Annotations'\r\n - Add a cleanup process to remove locked files\r\n - Make it SQL injection safe\r\n - Maybe make it look a bit nicer, it really is just a quick hack\r\n - Polish code and put it on github\r\n\r\n\r\n----------\r\nEdit:\r\n\r\n**Changes since initial post:**\r\n\r\n - Reset button actually does something now\r\n - Added a reference to libraries used\r\n\r\n  [1]: http://tobycheese.pythonanywhere.com/\r\n  [2]: http://tobycheese.pythonanywhere.com/annotated/img_72999\r\n  [3]: http://tobycheese.pythonanywhere.com/annotated/",
      "votes": 4
    },
    {
      "id": 127296,
      "postDate": "2016-07-13T20:08:59.030Z",
      "content": "<p>OK, thank to all who replied.</p>\n\n<p>I'll dig this back out once the next computer vision competition is launched where there are things to detect that are not already handled by a pretrained net (such as whales :-) )</p>",
      "rawMarkdown": "OK, thank to all who replied.\r\n\r\nI'll dig this back out once the next computer vision competition is launched where there are things to detect that are not already handled by a pretrained net (such as whales :-) )",
      "votes": 1
    },
    {
      "id": 117259,
      "postDate": "2016-04-28T08:04:42.217Z",
      "content": "<p>@Toby\nProbably a little bit of both. There is no clear benefits (it is relatively easy to get face and hands using simple image processing - but even that may be not needed - or even make things worse - for the CNN) and installing things could be a headache. </p>",
      "rawMarkdown": "@Toby\r\nProbably a little bit of both. There is no clear benefits (it is relatively easy to get face and hands using simple image processing - but even that may be not needed - or even make things worse - for the CNN) and installing things could be a headache. ",
      "votes": 1
    },
    {
      "id": 117145,
      "postDate": "2016-04-27T15:17:05.687Z",
      "content": "<p>Right, so now I even have to reply to my own posts, since nobody else does :-)</p>\n\n<p>Seriously though, I am a bit surprised that my attempt to start and guide a community effort attracted almost zero reaction so far, and I am wondering why that is or what I did wrong in the process.</p>\n\n<p>Is it simply because nobody needs an annotated training set in this competition?</p>\n\n<p>Or is the part where one has to install an extension for chrome too complicated, or too suspicious? This seems possible, because if the demo I put up was tried by anyone at all, no one has actually saved an annotated image so far. But then again, my personal opinion is that serving the image files from one&#8217;s local machine is actually a clever move (if I may say so). The images load and display instantly and there is no wasted bandwidth for downloading files that everyone already has downloaded.</p>\n\n<p>So, I am grateful for all feedback on why there is no feedback :-)</p>",
      "rawMarkdown": "Right, so now I even have to reply to my own posts, since nobody else does :-)\r\n\r\nSeriously though, I am a bit surprised that my attempt to start and guide a community effort attracted almost zero reaction so far, and I am wondering why that is or what I did wrong in the process.\r\n\r\nIs it simply because nobody needs an annotated training set in this competition?\r\n\r\nOr is the part where one has to install an extension for chrome too complicated, or too suspicious? This seems possible, because if the demo I put up was tried by anyone at all, no one has actually saved an annotated image so far. But then again, my personal opinion is that serving the image files from one’s local machine is actually a clever move (if I may say so). The images load and display instantly and there is no wasted bandwidth for downloading files that everyone already has downloaded.\r\n\r\nSo, I am grateful for all feedback on why there is no feedback :-)\r\n"
    },
    {
      "id": 117301,
      "postDate": "2016-04-28T12:23:43.727Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 117150,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-04-27T15:46:27.187000",
      "content": "<p>@Toby -</p>\n\n<p>I'm taking a guess here. Either (a) people think hand annotations will be too much work, and might not give incremental benefit over, e.g., pre-trained networks, or (b) people are annotating their images themselves to keep it as a competitive advantage.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 127296,
      "author_name": "Toby Cheese",
      "author_url": "",
      "post_date": "2016-07-13T20:08:59.030000",
      "content": "<p>OK, thank to all who replied.</p>\n\n<p>I'll dig this back out once the next computer vision competition is launched where there are things to detect that are not already handled by a pretrained net (such as whales :-) )</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 117259,
      "author_name": "==>(AL)<==",
      "author_url": "",
      "post_date": "2016-04-28T08:04:42.217000",
      "content": "<p>@Toby\nProbably a little bit of both. There is no clear benefits (it is relatively easy to get face and hands using simple image processing - but even that may be not needed - or even make things worse - for the CNN) and installing things could be a headache. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 117145,
      "author_name": "Toby Cheese",
      "author_url": "",
      "post_date": "2016-04-27T15:17:05.687000",
      "content": "<p>Right, so now I even have to reply to my own posts, since nobody else does :-)</p>\n\n<p>Seriously though, I am a bit surprised that my attempt to start and guide a community effort attracted almost zero reaction so far, and I am wondering why that is or what I did wrong in the process.</p>\n\n<p>Is it simply because nobody needs an annotated training set in this competition?</p>\n\n<p>Or is the part where one has to install an extension for chrome too complicated, or too suspicious? This seems possible, because if the demo I put up was tried by anyone at all, no one has actually saved an annotated image so far. But then again, my personal opinion is that serving the image files from one&#8217;s local machine is actually a clever move (if I may say so). The images load and display instantly and there is no wasted bandwidth for downloading files that everyone already has downloaded.</p>\n\n<p>So, I am grateful for all feedback on why there is no feedback :-)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 117301,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-04-28T12:23:43.727000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "117150": "@Toby -\r\n\r\nI'm taking a guess here. Either (a) people think hand annotations will be too much work, and might not give incremental benefit over, e.g., pre-trained networks, or (b) people are annotating their images themselves to keep it as a competitive advantage.",
    "116193": "Hi all\r\n\r\nI hacked together a web application so we can work on annotating the training set. You find the first version [here][1]. Before launching the effort for real, I need some feedback on a few things though:\r\n\r\n\r\n**What parts of the image should we annotate**?\r\nAs a start, I configured three ROIs per image: head, steering wheel and left arm. But we need a consensus as which parts we want to annotate before starting the effort.\r\n\r\n\r\n**In what order should the images appear for tagging?**\r\nAt the moment, the images are given in the order they are globbed from the filesystem. But this is suboptimal. One suggestion would be to bundle images per driver and in consequence to allow only as many people as there are distinct drivers to work in parallel. When doing it this way, for each new image the previously made annotations could be suggested, which would speed up annotations for a sequence of very similar images considerably. \r\n\r\n**How could handling be improved?**\r\nPlease play with it a little, and report ideas on how the annotation process can be more speedy or more convenient.\r\n\r\n**Bugs?**\r\nThey are somewhere, please report any sightings.\r\n\r\n**How should the (final) output be formatted?**\r\nThis is maybe less important, but at the moment output is simply all fields of the database dumped into json. If someone has a good suggestion, I can make output more convenient.\r\n\r\n\r\n**Some more notes:**\r\n\r\nTo access the results, use the following two links to either get the results of one single image, or of all images annotated so far:\r\n[http://tobycheese.pythonanywhere.com/annotated/img_72999][2]\r\n[http://tobycheese.pythonanywhere.com/annotated][3]\r\n\r\nIf you have already spent a substantial amount of work or money on annotating images and want to contribute this, we will find a way. At first I need all image ids that you have already annotated, so I can filter them out to avoid double work. Then (after we have decided on which ROIs to annotate), I can set up a second instance of the web application with just your images to annotate the missing ROIS (in case the ROIs you annotated differ from those here). Then your work can be integrated here.\r\n\r\n**Things still on the todo list:**\r\n\r\n - Get answers to above questions and implement the consequences\r\n - Add keyboard shortcuts for 'Next' and 'Reset Annotations'\r\n - Add a cleanup process to remove locked files\r\n - Make it SQL injection safe\r\n - Maybe make it look a bit nicer, it really is just a quick hack\r\n - Polish code and put it on github\r\n\r\n\r\n----------\r\nEdit:\r\n\r\n**Changes since initial post:**\r\n\r\n - Reset button actually does something now\r\n - Added a reference to libraries used\r\n\r\n  [1]: http://tobycheese.pythonanywhere.com/\r\n  [2]: http://tobycheese.pythonanywhere.com/annotated/img_72999\r\n  [3]: http://tobycheese.pythonanywhere.com/annotated/",
    "127296": "OK, thank to all who replied.\r\n\r\nI'll dig this back out once the next computer vision competition is launched where there are things to detect that are not already handled by a pretrained net (such as whales :-) )",
    "117259": "@Toby\r\nProbably a little bit of both. There is no clear benefits (it is relatively easy to get face and hands using simple image processing - but even that may be not needed - or even make things worse - for the CNN) and installing things could be a headache. ",
    "117145": "Right, so now I even have to reply to my own posts, since nobody else does :-)\r\n\r\nSeriously though, I am a bit surprised that my attempt to start and guide a community effort attracted almost zero reaction so far, and I am wondering why that is or what I did wrong in the process.\r\n\r\nIs it simply because nobody needs an annotated training set in this competition?\r\n\r\nOr is the part where one has to install an extension for chrome too complicated, or too suspicious? This seems possible, because if the demo I put up was tried by anyone at all, no one has actually saved an annotated image so far. But then again, my personal opinion is that serving the image files from one’s local machine is actually a clever move (if I may say so). The images load and display instantly and there is no wasted bandwidth for downloading files that everyone already has downloaded.\r\n\r\nSo, I am grateful for all feedback on why there is no feedback :-)\r\n",
    "117301": ""
  }
}