{
  "id": 20116,
  "title": "Easiest and hardest to predict drivers from training set",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20116",
  "author_name": "",
  "post_date": "2016-04-13T18:39:37.063Z",
  "votes": 3,
  "comment_count": 5,
  "views": 1233,
  "content": "<p>At least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.</p>\n\n<p>Easiest to predict, CV loss tends to match training loss and accuracy over 80% is &quot;p045&quot;.</p>\n\n<p>Hardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is &quot;p072&quot;. In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.</p>\n\n<p>Anyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.</p>",
  "messages": [
    {
      "id": "114795",
      "postDate": "04/13/2016 18:39:37",
      "content": "<p>At least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.</p>\n\n<p>Easiest to predict, CV loss tends to match training loss and accuracy over 80% is &quot;p045&quot;.</p>\n\n<p>Hardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is &quot;p072&quot;. In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.</p>\n\n<p>Anyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.</p>",
      "rawMarkdown": "At least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.\r\n\r\nEasiest to predict, CV loss tends to match training loss and accuracy over 80% is \"p045\".\r\n\r\nHardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is \"p072\". In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.\r\n\r\nAnyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.",
      "votes": null
    },
    {
      "id": "114803",
      "postDate": "04/13/2016 19:15:47",
      "content": "<p>I just finished writing my K-Fold cross-validation drivers-based script and I'm seeing exactly this. The following is the output of my script:</p>\n\n<pre><code>Loading data.\nStarting cross-validation.\n\nFold 1 of 10.\nTrain on 20000 samples, validate on 2424 samples\nEpoch 1/3\n20000/20000 [==============================] - 12s - loss: 0.7490 - val_loss: 2.5662\nEpoch 2/3\n20000/20000 [==============================] - 11s - loss: 0.2148 - val_loss: 3.4269\nEpoch 3/3\n20000/20000 [==============================] - 11s - loss: 0.1378 - val_loss: 3.4602\n\nFold 2 of 10.\nTrain on 19234 samples, validate on 3190 samples\nEpoch 1/3\n19234/19234 [==============================] - 11s - loss: 0.7637 - val_loss: 2.3823\nEpoch 2/3\n19234/19234 [==============================] - 11s - loss: 0.2359 - val_loss: 3.0116\nEpoch 3/3\n19234/19234 [==============================] - 11s - loss: 0.1484 - val_loss: 2.5999\n\nFold 3 of 10.\nTrain on 18769 samples, validate on 3655 samples\nEpoch 1/3\n18769/18769 [==============================] - 11s - loss: 0.7642 - val_loss: 2.4533\nEpoch 2/3\n18769/18769 [==============================] - 11s - loss: 0.2343 - val_loss: 2.3807\nEpoch 3/3\n18769/18769 [==============================] - 10s - loss: 0.1456 - val_loss: 2.4895\n\nFold 4 of 10.\nTrain on 20320 samples, validate on 2104 samples\nEpoch 1/3\n20320/20320 [==============================] - 11s - loss: 0.7503 - val_loss: 1.3538\nEpoch 2/3\n20320/20320 [==============================] - 11s - loss: 0.2151 - val_loss: 1.4942\nEpoch 3/3\n20320/20320 [==============================] - 11s - loss: 0.1391 - val_loss: 1.5492\n\nFold 5 of 10.\nTrain on 20274 samples, validate on 2150 samples\nEpoch 1/3\n20274/20274 [==============================] - 11s - loss: 0.7735 - val_loss: 1.1530\nEpoch 2/3\n20274/20274 [==============================] - 11s - loss: 0.2171 - val_loss: 1.2427\nEpoch 3/3\n20274/20274 [==============================] - 11s - loss: 0.1361 - val_loss: 1.1809\n\nFold 6 of 10.\nTrain on 19703 samples, validate on 2721 samples\nEpoch 1/3\n19703/19703 [==============================] - 11s - loss: 0.7533 - val_loss: 2.0541\nEpoch 2/3\n19703/19703 [==============================] - 11s - loss: 0.2069 - val_loss: 2.1339\nEpoch 3/3\n19703/19703 [==============================] - 11s - loss: 0.1445 - val_loss: 2.2169\n\nFold 7 of 10.\nTrain on 20890 samples, validate on 1534 samples\nEpoch 1/3\n20890/20890 [==============================] - 11s - loss: 0.7764 - val_loss: 0.9428\nEpoch 2/3\n20890/20890 [==============================] - 11s - loss: 0.2316 - val_loss: 0.8675\nEpoch 3/3\n20890/20890 [==============================] - 11s - loss: 0.1502 - val_loss: 0.8676\n\nFold 8 of 10.\nTrain on 20795 samples, validate on 1629 samples\nEpoch 1/3\n20795/20795 [==============================] - 11s - loss: 0.7433 - val_loss: 2.1116\nEpoch 2/3\n20795/20795 [==============================] - 11s - loss: 0.2177 - val_loss: 2.2820\nEpoch 3/3\n20795/20795 [==============================] - 11s - loss: 0.1428 - val_loss: 2.5316\n\nFold 9 of 10.\nTrain on 21044 samples, validate on 1380 samples\nEpoch 1/3\n21044/21044 [==============================] - 11s - loss: 0.7251 - val_loss: 2.3594\nEpoch 2/3\n21044/21044 [==============================] - 11s - loss: 0.2167 - val_loss: 2.3987\nEpoch 3/3\n21044/21044 [==============================] - 11s - loss: 0.1312 - val_loss: 2.6074\n\nFold 10 of 10.\nTrain on 20787 samples, validate on 1637 samples\nEpoch 1/3\n20787/20787 [==============================] - 12s - loss: 0.7307 - val_loss: 2.1836\nEpoch 2/3\n20787/20787 [==============================] - 11s - loss: 0.2158 - val_loss: 2.2413\nEpoch 3/3\n20787/20787 [==============================] - 11s - loss: 0.1457 - val_loss: 2.4234\n\nResults:\nEpoch | Loss\n1     | 1.95601125879 +- 1.11063694572\n2     | 2.14794619592 +- 1.47128580809\n3     | 2.19266234801 +- 1.46898164285\n</code></pre>\n\n<p>It's possible to see that there were some folds, such as 1, 9 and 10, where the validation loss was pretty high, while in fold 5 and 7 it was very small. This could suggest that the drivers in the folds it had a bad performance are harder to predict.</p>\n\n<p>Also, it's important to notice that I haven't tuned this model yet, so maybe this wouldn't happen for a tuned model.</p>",
      "rawMarkdown": "I just finished writing my K-Fold cross-validation drivers-based script and I'm seeing exactly this. The following is the output of my script:\r\n\r\n    Loading data.\r\n    Starting cross-validation.\r\n    \r\n    Fold 1 of 10.\r\n    Train on 20000 samples, validate on 2424 samples\r\n    Epoch 1/3\r\n    20000/20000 [==============================] - 12s - loss: 0.7490 - val_loss: 2.5662\r\n    Epoch 2/3\r\n    20000/20000 [==============================] - 11s - loss: 0.2148 - val_loss: 3.4269\r\n    Epoch 3/3\r\n    20000/20000 [==============================] - 11s - loss: 0.1378 - val_loss: 3.4602\r\n    \r\n    Fold 2 of 10.\r\n    Train on 19234 samples, validate on 3190 samples\r\n    Epoch 1/3\r\n    19234/19234 [==============================] - 11s - loss: 0.7637 - val_loss: 2.3823\r\n    Epoch 2/3\r\n    19234/19234 [==============================] - 11s - loss: 0.2359 - val_loss: 3.0116\r\n    Epoch 3/3\r\n    19234/19234 [==============================] - 11s - loss: 0.1484 - val_loss: 2.5999\r\n    \r\n    Fold 3 of 10.\r\n    Train on 18769 samples, validate on 3655 samples\r\n    Epoch 1/3\r\n    18769/18769 [==============================] - 11s - loss: 0.7642 - val_loss: 2.4533\r\n    Epoch 2/3\r\n    18769/18769 [==============================] - 11s - loss: 0.2343 - val_loss: 2.3807\r\n    Epoch 3/3\r\n    18769/18769 [==============================] - 10s - loss: 0.1456 - val_loss: 2.4895\r\n    \r\n    Fold 4 of 10.\r\n    Train on 20320 samples, validate on 2104 samples\r\n    Epoch 1/3\r\n    20320/20320 [==============================] - 11s - loss: 0.7503 - val_loss: 1.3538\r\n    Epoch 2/3\r\n    20320/20320 [==============================] - 11s - loss: 0.2151 - val_loss: 1.4942\r\n    Epoch 3/3\r\n    20320/20320 [==============================] - 11s - loss: 0.1391 - val_loss: 1.5492\r\n    \r\n    Fold 5 of 10.\r\n    Train on 20274 samples, validate on 2150 samples\r\n    Epoch 1/3\r\n    20274/20274 [==============================] - 11s - loss: 0.7735 - val_loss: 1.1530\r\n    Epoch 2/3\r\n    20274/20274 [==============================] - 11s - loss: 0.2171 - val_loss: 1.2427\r\n    Epoch 3/3\r\n    20274/20274 [==============================] - 11s - loss: 0.1361 - val_loss: 1.1809\r\n    \r\n    Fold 6 of 10.\r\n    Train on 19703 samples, validate on 2721 samples\r\n    Epoch 1/3\r\n    19703/19703 [==============================] - 11s - loss: 0.7533 - val_loss: 2.0541\r\n    Epoch 2/3\r\n    19703/19703 [==============================] - 11s - loss: 0.2069 - val_loss: 2.1339\r\n    Epoch 3/3\r\n    19703/19703 [==============================] - 11s - loss: 0.1445 - val_loss: 2.2169\r\n    \r\n    Fold 7 of 10.\r\n    Train on 20890 samples, validate on 1534 samples\r\n    Epoch 1/3\r\n    20890/20890 [==============================] - 11s - loss: 0.7764 - val_loss: 0.9428\r\n    Epoch 2/3\r\n    20890/20890 [==============================] - 11s - loss: 0.2316 - val_loss: 0.8675\r\n    Epoch 3/3\r\n    20890/20890 [==============================] - 11s - loss: 0.1502 - val_loss: 0.8676\r\n    \r\n    Fold 8 of 10.\r\n    Train on 20795 samples, validate on 1629 samples\r\n    Epoch 1/3\r\n    20795/20795 [==============================] - 11s - loss: 0.7433 - val_loss: 2.1116\r\n    Epoch 2/3\r\n    20795/20795 [==============================] - 11s - loss: 0.2177 - val_loss: 2.2820\r\n    Epoch 3/3\r\n    20795/20795 [==============================] - 11s - loss: 0.1428 - val_loss: 2.5316\r\n    \r\n    Fold 9 of 10.\r\n    Train on 21044 samples, validate on 1380 samples\r\n    Epoch 1/3\r\n    21044/21044 [==============================] - 11s - loss: 0.7251 - val_loss: 2.3594\r\n    Epoch 2/3\r\n    21044/21044 [==============================] - 11s - loss: 0.2167 - val_loss: 2.3987\r\n    Epoch 3/3\r\n    21044/21044 [==============================] - 11s - loss: 0.1312 - val_loss: 2.6074\r\n    \r\n    Fold 10 of 10.\r\n    Train on 20787 samples, validate on 1637 samples\r\n    Epoch 1/3\r\n    20787/20787 [==============================] - 12s - loss: 0.7307 - val_loss: 2.1836\r\n    Epoch 2/3\r\n    20787/20787 [==============================] - 11s - loss: 0.2158 - val_loss: 2.2413\r\n    Epoch 3/3\r\n    20787/20787 [==============================] - 11s - loss: 0.1457 - val_loss: 2.4234\r\n    \r\n    Results:\r\n    Epoch | Loss\r\n    1     | 1.95601125879 +- 1.11063694572\r\n    2     | 2.14794619592 +- 1.47128580809\r\n    3     | 2.19266234801 +- 1.46898164285\r\n\r\nIt's possible to see that there were some folds, such as 1, 9 and 10, where the validation loss was pretty high, while in fold 5 and 7 it was very small. This could suggest that the drivers in the folds it had a bad performance are harder to predict.\r\n\r\nAlso, it's important to notice that I haven't tuned this model yet, so maybe this wouldn't happen for a tuned model.",
      "votes": null
    },
    {
      "id": "114930",
      "postDate": "04/14/2016 20:04:10",
      "content": "<p>I have very bad CV result for both P072 and P050.  </p>\n\n<p>[quote=Neil Slater;114795]</p>\n\n<p>At least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.</p>\n\n<p>Easiest to predict, CV loss tends to match training loss and accuracy over 80% is &quot;p045&quot;.</p>\n\n<p>Hardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is &quot;p072&quot;. In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.</p>\n\n<p>Anyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I have very bad CV result for both P072 and P050.  \r\n\r\n\r\n[quote=Neil Slater;114795]\r\n\r\nAt least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.\r\n\r\nEasiest to predict, CV loss tends to match training loss and accuracy over 80% is \"p045\".\r\n\r\nHardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is \"p072\". In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.\r\n\r\nAnyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.\r\n\r\n[/quote]",
      "votes": null
    },
    {
      "id": "115055",
      "postDate": "04/15/2016 23:06:10",
      "content": "<p>No surprises here.</p>\n\n<p>p050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \nso the network fail to identify arms - one of the strongest differentiators. </p>\n\n<p>p072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set. </p>",
      "rawMarkdown": "No surprises here.\r\n\r\np050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \r\nso the network fail to identify arms - one of the strongest differentiators. \r\n\r\np072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set.",
      "votes": null
    },
    {
      "id": "115058",
      "postDate": "04/15/2016 23:26:23",
      "content": "<p>How about error analysis on predicted classes. Does anybody have a clue?</p>",
      "rawMarkdown": "How about error analysis on predicted classes. Does anybody have a clue?",
      "votes": null
    },
    {
      "id": "115096",
      "postDate": "04/16/2016 07:30:25",
      "content": "<p>[quote===&gt;(AL)&lt;==;115055]</p>\n\n<p>No surprises here.</p>\n\n<p>p050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \nso the network fail to identify arms - one of the strongest differentiators. </p>\n\n<p>p072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set. </p>\n\n<p>[/quote]</p>\n\n<p>Is there any manipulations we could do to existing drivers that might stop these being outliers? It should definitely be within realms of possibility to alter colour of clothing, although I am not sure I could find an automated filter accurate and reliable enough to apply to large numbers of images. More likely the network would learn to spot artefacts from the manipulations. </p>\n\n<p>Perhaps now we are allowed external data, a network pre-trained for human outline segmentation on a larger number and variety of people might help?</p>",
      "rawMarkdown": "[quote===>(AL)<==;115055]\r\n\r\nNo surprises here.\r\n\r\np050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \r\nso the network fail to identify arms - one of the strongest differentiators. \r\n\r\np072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set. \r\n\r\n[/quote]\r\n\r\nIs there any manipulations we could do to existing drivers that might stop these being outliers? It should definitely be within realms of possibility to alter colour of clothing, although I am not sure I could find an automated filter accurate and reliable enough to apply to large numbers of images. More likely the network would learn to spot artefacts from the manipulations. \r\n\r\nPerhaps now we are allowed external data, a network pre-trained for human outline segmentation on a larger number and variety of people might help?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 114803,
      "author_name": "pennacchio",
      "author_url": "",
      "post_date": "04/13/2016 19:15:47",
      "content": "<p>I just finished writing my K-Fold cross-validation drivers-based script and I'm seeing exactly this. The following is the output of my script:</p>\n\n<pre><code>Loading data.\nStarting cross-validation.\n\nFold 1 of 10.\nTrain on 20000 samples, validate on 2424 samples\nEpoch 1/3\n20000/20000 [==============================] - 12s - loss: 0.7490 - val_loss: 2.5662\nEpoch 2/3\n20000/20000 [==============================] - 11s - loss: 0.2148 - val_loss: 3.4269\nEpoch 3/3\n20000/20000 [==============================] - 11s - loss: 0.1378 - val_loss: 3.4602\n\nFold 2 of 10.\nTrain on 19234 samples, validate on 3190 samples\nEpoch 1/3\n19234/19234 [==============================] - 11s - loss: 0.7637 - val_loss: 2.3823\nEpoch 2/3\n19234/19234 [==============================] - 11s - loss: 0.2359 - val_loss: 3.0116\nEpoch 3/3\n19234/19234 [==============================] - 11s - loss: 0.1484 - val_loss: 2.5999\n\nFold 3 of 10.\nTrain on 18769 samples, validate on 3655 samples\nEpoch 1/3\n18769/18769 [==============================] - 11s - loss: 0.7642 - val_loss: 2.4533\nEpoch 2/3\n18769/18769 [==============================] - 11s - loss: 0.2343 - val_loss: 2.3807\nEpoch 3/3\n18769/18769 [==============================] - 10s - loss: 0.1456 - val_loss: 2.4895\n\nFold 4 of 10.\nTrain on 20320 samples, validate on 2104 samples\nEpoch 1/3\n20320/20320 [==============================] - 11s - loss: 0.7503 - val_loss: 1.3538\nEpoch 2/3\n20320/20320 [==============================] - 11s - loss: 0.2151 - val_loss: 1.4942\nEpoch 3/3\n20320/20320 [==============================] - 11s - loss: 0.1391 - val_loss: 1.5492\n\nFold 5 of 10.\nTrain on 20274 samples, validate on 2150 samples\nEpoch 1/3\n20274/20274 [==============================] - 11s - loss: 0.7735 - val_loss: 1.1530\nEpoch 2/3\n20274/20274 [==============================] - 11s - loss: 0.2171 - val_loss: 1.2427\nEpoch 3/3\n20274/20274 [==============================] - 11s - loss: 0.1361 - val_loss: 1.1809\n\nFold 6 of 10.\nTrain on 19703 samples, validate on 2721 samples\nEpoch 1/3\n19703/19703 [==============================] - 11s - loss: 0.7533 - val_loss: 2.0541\nEpoch 2/3\n19703/19703 [==============================] - 11s - loss: 0.2069 - val_loss: 2.1339\nEpoch 3/3\n19703/19703 [==============================] - 11s - loss: 0.1445 - val_loss: 2.2169\n\nFold 7 of 10.\nTrain on 20890 samples, validate on 1534 samples\nEpoch 1/3\n20890/20890 [==============================] - 11s - loss: 0.7764 - val_loss: 0.9428\nEpoch 2/3\n20890/20890 [==============================] - 11s - loss: 0.2316 - val_loss: 0.8675\nEpoch 3/3\n20890/20890 [==============================] - 11s - loss: 0.1502 - val_loss: 0.8676\n\nFold 8 of 10.\nTrain on 20795 samples, validate on 1629 samples\nEpoch 1/3\n20795/20795 [==============================] - 11s - loss: 0.7433 - val_loss: 2.1116\nEpoch 2/3\n20795/20795 [==============================] - 11s - loss: 0.2177 - val_loss: 2.2820\nEpoch 3/3\n20795/20795 [==============================] - 11s - loss: 0.1428 - val_loss: 2.5316\n\nFold 9 of 10.\nTrain on 21044 samples, validate on 1380 samples\nEpoch 1/3\n21044/21044 [==============================] - 11s - loss: 0.7251 - val_loss: 2.3594\nEpoch 2/3\n21044/21044 [==============================] - 11s - loss: 0.2167 - val_loss: 2.3987\nEpoch 3/3\n21044/21044 [==============================] - 11s - loss: 0.1312 - val_loss: 2.6074\n\nFold 10 of 10.\nTrain on 20787 samples, validate on 1637 samples\nEpoch 1/3\n20787/20787 [==============================] - 12s - loss: 0.7307 - val_loss: 2.1836\nEpoch 2/3\n20787/20787 [==============================] - 11s - loss: 0.2158 - val_loss: 2.2413\nEpoch 3/3\n20787/20787 [==============================] - 11s - loss: 0.1457 - val_loss: 2.4234\n\nResults:\nEpoch | Loss\n1     | 1.95601125879 +- 1.11063694572\n2     | 2.14794619592 +- 1.47128580809\n3     | 2.19266234801 +- 1.46898164285\n</code></pre>\n\n<p>It's possible to see that there were some folds, such as 1, 9 and 10, where the validation loss was pretty high, while in fold 5 and 7 it was very small. This could suggest that the drivers in the folds it had a bad performance are harder to predict.</p>\n\n<p>Also, it's important to notice that I haven't tuned this model yet, so maybe this wouldn't happen for a tuned model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 114930,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "04/14/2016 20:04:10",
      "content": "<p>I have very bad CV result for both P072 and P050.  </p>\n\n<p>[quote=Neil Slater;114795]</p>\n\n<p>At least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.</p>\n\n<p>Easiest to predict, CV loss tends to match training loss and accuracy over 80% is &quot;p045&quot;.</p>\n\n<p>Hardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is &quot;p072&quot;. In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.</p>\n\n<p>Anyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.</p>\n\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115055,
      "author_name": "alexlzzz",
      "author_url": "",
      "post_date": "04/15/2016 23:06:10",
      "content": "<p>No surprises here.</p>\n\n<p>p050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \nso the network fail to identify arms - one of the strongest differentiators. </p>\n\n<p>p072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115058,
      "author_name": "potamitis",
      "author_url": "",
      "post_date": "04/15/2016 23:26:23",
      "content": "<p>How about error analysis on predicted classes. Does anybody have a clue?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 115096,
      "author_name": "slobo777",
      "author_url": "",
      "post_date": "04/16/2016 07:30:25",
      "content": "<p>[quote===&gt;(AL)&lt;==;115055]</p>\n\n<p>No surprises here.</p>\n\n<p>p050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \nso the network fail to identify arms - one of the strongest differentiators. </p>\n\n<p>p072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set. </p>\n\n<p>[/quote]</p>\n\n<p>Is there any manipulations we could do to existing drivers that might stop these being outliers? It should definitely be within realms of possibility to alter colour of clothing, although I am not sure I could find an automated filter accurate and reliable enough to apply to large numbers of images. More likely the network would learn to spot artefacts from the manipulations. </p>\n\n<p>Perhaps now we are allowed external data, a network pre-trained for human outline segmentation on a larger number and variety of people might help?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "114795": "At least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.\r\n\r\nEasiest to predict, CV loss tends to match training loss and accuracy over 80% is \"p045\".\r\n\r\nHardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is \"p072\". In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.\r\n\r\nAnyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.",
    "114803": "I just finished writing my K-Fold cross-validation drivers-based script and I'm seeing exactly this. The following is the output of my script:\r\n\r\n    Loading data.\r\n    Starting cross-validation.\r\n    \r\n    Fold 1 of 10.\r\n    Train on 20000 samples, validate on 2424 samples\r\n    Epoch 1/3\r\n    20000/20000 [==============================] - 12s - loss: 0.7490 - val_loss: 2.5662\r\n    Epoch 2/3\r\n    20000/20000 [==============================] - 11s - loss: 0.2148 - val_loss: 3.4269\r\n    Epoch 3/3\r\n    20000/20000 [==============================] - 11s - loss: 0.1378 - val_loss: 3.4602\r\n    \r\n    Fold 2 of 10.\r\n    Train on 19234 samples, validate on 3190 samples\r\n    Epoch 1/3\r\n    19234/19234 [==============================] - 11s - loss: 0.7637 - val_loss: 2.3823\r\n    Epoch 2/3\r\n    19234/19234 [==============================] - 11s - loss: 0.2359 - val_loss: 3.0116\r\n    Epoch 3/3\r\n    19234/19234 [==============================] - 11s - loss: 0.1484 - val_loss: 2.5999\r\n    \r\n    Fold 3 of 10.\r\n    Train on 18769 samples, validate on 3655 samples\r\n    Epoch 1/3\r\n    18769/18769 [==============================] - 11s - loss: 0.7642 - val_loss: 2.4533\r\n    Epoch 2/3\r\n    18769/18769 [==============================] - 11s - loss: 0.2343 - val_loss: 2.3807\r\n    Epoch 3/3\r\n    18769/18769 [==============================] - 10s - loss: 0.1456 - val_loss: 2.4895\r\n    \r\n    Fold 4 of 10.\r\n    Train on 20320 samples, validate on 2104 samples\r\n    Epoch 1/3\r\n    20320/20320 [==============================] - 11s - loss: 0.7503 - val_loss: 1.3538\r\n    Epoch 2/3\r\n    20320/20320 [==============================] - 11s - loss: 0.2151 - val_loss: 1.4942\r\n    Epoch 3/3\r\n    20320/20320 [==============================] - 11s - loss: 0.1391 - val_loss: 1.5492\r\n    \r\n    Fold 5 of 10.\r\n    Train on 20274 samples, validate on 2150 samples\r\n    Epoch 1/3\r\n    20274/20274 [==============================] - 11s - loss: 0.7735 - val_loss: 1.1530\r\n    Epoch 2/3\r\n    20274/20274 [==============================] - 11s - loss: 0.2171 - val_loss: 1.2427\r\n    Epoch 3/3\r\n    20274/20274 [==============================] - 11s - loss: 0.1361 - val_loss: 1.1809\r\n    \r\n    Fold 6 of 10.\r\n    Train on 19703 samples, validate on 2721 samples\r\n    Epoch 1/3\r\n    19703/19703 [==============================] - 11s - loss: 0.7533 - val_loss: 2.0541\r\n    Epoch 2/3\r\n    19703/19703 [==============================] - 11s - loss: 0.2069 - val_loss: 2.1339\r\n    Epoch 3/3\r\n    19703/19703 [==============================] - 11s - loss: 0.1445 - val_loss: 2.2169\r\n    \r\n    Fold 7 of 10.\r\n    Train on 20890 samples, validate on 1534 samples\r\n    Epoch 1/3\r\n    20890/20890 [==============================] - 11s - loss: 0.7764 - val_loss: 0.9428\r\n    Epoch 2/3\r\n    20890/20890 [==============================] - 11s - loss: 0.2316 - val_loss: 0.8675\r\n    Epoch 3/3\r\n    20890/20890 [==============================] - 11s - loss: 0.1502 - val_loss: 0.8676\r\n    \r\n    Fold 8 of 10.\r\n    Train on 20795 samples, validate on 1629 samples\r\n    Epoch 1/3\r\n    20795/20795 [==============================] - 11s - loss: 0.7433 - val_loss: 2.1116\r\n    Epoch 2/3\r\n    20795/20795 [==============================] - 11s - loss: 0.2177 - val_loss: 2.2820\r\n    Epoch 3/3\r\n    20795/20795 [==============================] - 11s - loss: 0.1428 - val_loss: 2.5316\r\n    \r\n    Fold 9 of 10.\r\n    Train on 21044 samples, validate on 1380 samples\r\n    Epoch 1/3\r\n    21044/21044 [==============================] - 11s - loss: 0.7251 - val_loss: 2.3594\r\n    Epoch 2/3\r\n    21044/21044 [==============================] - 11s - loss: 0.2167 - val_loss: 2.3987\r\n    Epoch 3/3\r\n    21044/21044 [==============================] - 11s - loss: 0.1312 - val_loss: 2.6074\r\n    \r\n    Fold 10 of 10.\r\n    Train on 20787 samples, validate on 1637 samples\r\n    Epoch 1/3\r\n    20787/20787 [==============================] - 12s - loss: 0.7307 - val_loss: 2.1836\r\n    Epoch 2/3\r\n    20787/20787 [==============================] - 11s - loss: 0.2158 - val_loss: 2.2413\r\n    Epoch 3/3\r\n    20787/20787 [==============================] - 11s - loss: 0.1457 - val_loss: 2.4234\r\n    \r\n    Results:\r\n    Epoch | Loss\r\n    1     | 1.95601125879 +- 1.11063694572\r\n    2     | 2.14794619592 +- 1.47128580809\r\n    3     | 2.19266234801 +- 1.46898164285\r\n\r\nIt's possible to see that there were some folds, such as 1, 9 and 10, where the validation loss was pretty high, while in fold 5 and 7 it was very small. This could suggest that the drivers in the folds it had a bad performance are harder to predict.\r\n\r\nAlso, it's important to notice that I haven't tuned this model yet, so maybe this wouldn't happen for a tuned model.",
    "114930": "I have very bad CV result for both P072 and P050.  \r\n\r\n\r\n[quote=Neil Slater;114795]\r\n\r\nAt least with my approach so far - which is basically hacking and tuning ZFTurbo's excellent starter script - I have consistent easy-to-predict and hard-to-predict drivers when they are in CV set.\r\n\r\nEasiest to predict, CV loss tends to match training loss and accuracy over 80% is \"p045\".\r\n\r\nHardest to predict with CV loss and accuracy so bad (consistently worse than 10% - i.e. guessing - on my best model so far) I suspect something seriously amiss is \"p072\". In fact as the model trains, the accuracy goes down to 3% for this driver, and loss over 7! This implies some very confident mis-categorisation from my model. I'm planning to work through a few sample images this evening to see if I can find what is causing this effect.\r\n\r\nAnyone else seeing the same, or have other candidates? It would be interesting to know if it is some assumptions in my model causing this preferred driver effect, or if it is inherent to the data.\r\n\r\n[/quote]",
    "115055": "No surprises here.\r\n\r\np050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \r\nso the network fail to identify arms - one of the strongest differentiators. \r\n\r\np072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set.",
    "115058": "How about error analysis on predicted classes. Does anybody have a clue?",
    "115096": "[quote===>(AL)<==;115055]\r\n\r\nNo surprises here.\r\n\r\np050 is a dark skinned lady dressed with long black sleeves, about the same color as the cars, covering entire arms to the fingers. \r\nso the network fail to identify arms - one of the strongest differentiators. \r\n\r\np072 is a very large lady, probably distorting geometry network learned based on all other subjects in the train set. \r\n\r\n[/quote]\r\n\r\nIs there any manipulations we could do to existing drivers that might stop these being outliers? It should definitely be within realms of possibility to alter colour of clothing, although I am not sure I could find an automated filter accurate and reliable enough to apply to large numbers of images. More likely the network would learn to spot artefacts from the manipulations. \r\n\r\nPerhaps now we are allowed external data, a network pre-trained for human outline segmentation on a larger number and variety of people might help?"
  },
  "source": "meta"
}