{
  "id": 22298,
  "title": "Can the test set images be processed in batch mode",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22298",
  "author_name": "",
  "post_date": "2016-07-15T16:07:32.560Z",
  "votes": null,
  "comment_count": 8,
  "views": 1007,
  "content": "<p>If I would use some kind of automated method to separate the test set into different driver sequences and then predict the labels for those sequences, would it be legal? This wouldn't consist of any manual labeling and it would also work for any new set of images (that would also have a similar structure - frames extracted from different driver sequences), so it seems it should be allowed, but I'm just not sure about what is allowed to do with the testing set. Can anyone clarify?</p>",
  "messages": [
    {
      "id": "127827",
      "postDate": "07/15/2016 16:07:32",
      "content": "<p>If I would use some kind of automated method to separate the test set into different driver sequences and then predict the labels for those sequences, would it be legal? This wouldn't consist of any manual labeling and it would also work for any new set of images (that would also have a similar structure - frames extracted from different driver sequences), so it seems it should be allowed, but I'm just not sure about what is allowed to do with the testing set. Can anyone clarify?</p>",
      "rawMarkdown": "If I would use some kind of automated method to separate the test set into different driver sequences and then predict the labels for those sequences, would it be legal? This wouldn't consist of any manual labeling and it would also work for any new set of images (that would also have a similar structure - frames extracted from different driver sequences), so it seems it should be allowed, but I'm just not sure about what is allowed to do with the testing set. Can anyone clarify?",
      "votes": null
    },
    {
      "id": "128027",
      "postDate": "07/17/2016 05:46:27",
      "content": "<p>Hi bobutis,</p>\n\n<p>Your idea is interesting.<br>\nI also want to know whether grouping test images is allowed.</p>",
      "rawMarkdown": "Hi bobutis,\r\n\r\nYour idea is interesting.<br>\r\nI also want to know whether grouping test images is allowed.",
      "votes": null
    },
    {
      "id": "128070",
      "postDate": "07/17/2016 14:02:47",
      "content": "<p>bump</p>",
      "rawMarkdown": "bump",
      "votes": null
    },
    {
      "id": "128575",
      "postDate": "07/21/2016 09:56:39",
      "content": "<p>Can we get an answer from Competition Admin here please? My last submit does process test images in batch mode (which lead to high improvement and I'm sure much more improvement is possible) so this is important. Would a top3 solution be accepted if it does it? </p>",
      "rawMarkdown": "Can we get an answer from Competition Admin here please? My last submit does process test images in batch mode (which lead to high improvement and I'm sure much more improvement is possible) so this is important. Would a top3 solution be accepted if it does it?",
      "votes": null
    },
    {
      "id": "128936",
      "postDate": "07/25/2016 09:50:18",
      "content": "<p>The more generalized question is probably... are semi-supervised learning methods allowed? I haven't done it myself but I've always suspected running a convolutional autoencoder on the test set before training could give very good results.</p>",
      "rawMarkdown": "The more generalized question is probably... are semi-supervised learning methods allowed? I haven't done it myself but I've always suspected running a convolutional autoencoder on the test set before training could give very good results.",
      "votes": null
    },
    {
      "id": "128953",
      "postDate": "07/25/2016 12:20:55",
      "content": "<p>Usually semi-supervised methods are allowed, and there is nothing in the rules of this specific competition proscribing them.</p>",
      "rawMarkdown": "Usually semi-supervised methods are allowed, and there is nothing in the rules of this specific competition proscribing them.",
      "votes": null
    },
    {
      "id": "129009",
      "postDate": "07/25/2016 20:02:47",
      "content": "<p>Semi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine. </p>",
      "rawMarkdown": "Semi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine.",
      "votes": null
    },
    {
      "id": "129017",
      "postDate": "07/25/2016 20:59:57",
      "content": "<p>@Wendy Kan, as I understood the question (and I have the same question): is it legal to somehow find (with some preprocessing, not manual labelling of course) to which driver each of the test images belongs, and then use this information, for example, build driver-wise features. </p>\n\n<p>Currently driver-wise features would work on train and not on test. In real life, I think StateFarm has the information that all particular photos belong to the same driver.</p>\n\n<p>But if they want to score photo by photo and not the whole batch of photos then they might  get some of the features wrong. </p>\n\n<p>UPD: from the rules point of view this seems to be perfectly fine, and perhaps some people are already doing this</p>",
      "rawMarkdown": "Wendy Kan, as I understood the question (and I have the same question): is it legal to somehow find (with some preprocessing, not manual labelling of course) to which driver each of the test images belongs, and then use this information, for example, build driver-wise features. \r\n\r\nCurrently driver-wise features would work on train and not on test. In real life, I think StateFarm has the information that all particular photos belong to the same driver.\r\n\r\nBut if they want to score photo by photo and not the whole batch of photos then they might  get some of the features wrong. \r\n\r\nUPD: from the rules point of view this seems to be perfectly fine, and perhaps some people are already doing this",
      "votes": null
    },
    {
      "id": "129030",
      "postDate": "07/26/2016 01:05:32",
      "content": "<p>[quote=Wendy Kan;129009]</p>\n\n<p>Semi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine. </p>\n\n<p>[/quote]</p>\n\n<p>Of-couse, such model will have some ability to generalize. More data - the better generalization. But also you should realize that using test data (79k images) allows to enchance low/middle level feature extraction for that specific drivers from your test set. So, it will be a sort of overfitting. </p>",
      "rawMarkdown": "[quote=Wendy Kan;129009]\r\n\r\nSemi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine. \r\n\r\n[/quote]\r\n\r\nOf-couse, such model will have some ability to generalize. More data - the better generalization. But also you should realize that using test data (79k images) allows to enchance low/middle level feature extraction for that specific drivers from your test set. So, it will be a sort of overfitting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 128027,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "07/17/2016 05:46:27",
      "content": "<p>Hi bobutis,</p>\n\n<p>Your idea is interesting.<br>\nI also want to know whether grouping test images is allowed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 128070,
      "author_name": "assafler",
      "author_url": "",
      "post_date": "07/17/2016 14:02:47",
      "content": "<p>bump</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 128575,
      "author_name": "bobutis",
      "author_url": "",
      "post_date": "07/21/2016 09:56:39",
      "content": "<p>Can we get an answer from Competition Admin here please? My last submit does process test images in batch mode (which lead to high improvement and I'm sure much more improvement is possible) so this is important. Would a top3 solution be accepted if it does it? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 128936,
      "author_name": "martinkou",
      "author_url": "",
      "post_date": "07/25/2016 09:50:18",
      "content": "<p>The more generalized question is probably... are semi-supervised learning methods allowed? I haven't done it myself but I've always suspected running a convolutional autoencoder on the test set before training could give very good results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 128953,
      "author_name": "jadielam",
      "author_url": "",
      "post_date": "07/25/2016 12:20:55",
      "content": "<p>Usually semi-supervised methods are allowed, and there is nothing in the rules of this specific competition proscribing them.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129009,
      "author_name": "wendykan",
      "author_url": "",
      "post_date": "07/25/2016 20:02:47",
      "content": "<p>Semi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129017,
      "author_name": "vykhand",
      "author_url": "",
      "post_date": "07/25/2016 20:59:57",
      "content": "<p>@Wendy Kan, as I understood the question (and I have the same question): is it legal to somehow find (with some preprocessing, not manual labelling of course) to which driver each of the test images belongs, and then use this information, for example, build driver-wise features. </p>\n\n<p>Currently driver-wise features would work on train and not on test. In real life, I think StateFarm has the information that all particular photos belong to the same driver.</p>\n\n<p>But if they want to score photo by photo and not the whole batch of photos then they might  get some of the features wrong. </p>\n\n<p>UPD: from the rules point of view this seems to be perfectly fine, and perhaps some people are already doing this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129030,
      "author_name": "vasyutka",
      "author_url": "",
      "post_date": "07/26/2016 01:05:32",
      "content": "<p>[quote=Wendy Kan;129009]</p>\n\n<p>Semi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine. </p>\n\n<p>[/quote]</p>\n\n<p>Of-couse, such model will have some ability to generalize. More data - the better generalization. But also you should realize that using test data (79k images) allows to enchance low/middle level feature extraction for that specific drivers from your test set. So, it will be a sort of overfitting. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "127827": "If I would use some kind of automated method to separate the test set into different driver sequences and then predict the labels for those sequences, would it be legal? This wouldn't consist of any manual labeling and it would also work for any new set of images (that would also have a similar structure - frames extracted from different driver sequences), so it seems it should be allowed, but I'm just not sure about what is allowed to do with the testing set. Can anyone clarify?",
    "128027": "Hi bobutis,\r\n\r\nYour idea is interesting.<br>\r\nI also want to know whether grouping test images is allowed.",
    "128070": "bump",
    "128575": "Can we get an answer from Competition Admin here please? My last submit does process test images in batch mode (which lead to high improvement and I'm sure much more improvement is possible) so this is important. Would a top3 solution be accepted if it does it?",
    "128936": "The more generalized question is probably... are semi-supervised learning methods allowed? I haven't done it myself but I've always suspected running a convolutional autoencoder on the test set before training could give very good results.",
    "128953": "Usually semi-supervised methods are allowed, and there is nothing in the rules of this specific competition proscribing them.",
    "129009": "Semi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine.",
    "129017": "Wendy Kan, as I understood the question (and I have the same question): is it legal to somehow find (with some preprocessing, not manual labelling of course) to which driver each of the test images belongs, and then use this information, for example, build driver-wise features. \r\n\r\nCurrently driver-wise features would work on train and not on test. In real life, I think StateFarm has the information that all particular photos belong to the same driver.\r\n\r\nBut if they want to score photo by photo and not the whole batch of photos then they might  get some of the features wrong. \r\n\r\nUPD: from the rules point of view this seems to be perfectly fine, and perhaps some people are already doing this",
    "129030": "[quote=Wendy Kan;129009]\r\n\r\nSemi-supervised learning is allowed. The rule of thumb is, if you're given a new set of images tomorrow (new drivers, in this example), will the model be able to generalize? Hard to tell the details from @bobutis's description, but if the answer is yes, it's fine. \r\n\r\n[/quote]\r\n\r\nOf-couse, such model will have some ability to generalize. More data - the better generalization. But also you should realize that using test data (79k images) allows to enchance low/middle level feature extraction for that specific drivers from your test set. So, it will be a sort of overfitting."
  },
  "source": "meta"
}