{
  "id": 20775,
  "title": "Why is my leaderboard score so bad?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20775",
  "author_name": "ckleban",
  "post_date": "2016-05-06T18:46:42.897000",
  "votes": 1,
  "comment_count": 12,
  "views": 2342,
  "content": "<p>Hello,</p>\n\n<p>I'm new to Kaggle and ML and I'm running to an issue that I don't understand. </p>\n\n<p>When I submit the 'example submission' what has all 0.1 values, I get a score of about 2.3. When I train a simple DNN using tensorflow and skflow, I get a Test Accuracy of 0.97681 when comparing my trained model on my CV data. However, my LB score for this DNN is higher than the sample submission, at around 7 or higher. I'm not sure what could be causing this. My thoughts are:</p>\n\n<p>1) Maybe my submissions are wrong somehow? I'm currently submitting the probabilities, versus the predictions. Is that what I'm suppose to do? Fwiw, I'm using this function: classifier.predict_proba(X_test) to generate a 2D array of probabilities for the 70k+ test images. A small excerpt of my submission file looks like: </p>\n\n<p>img,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9\nimg_1.jpg,0.00000,0.00000,0.00034,0.00000,0.00000,0.99953,0.00013,0.00000,0.00000,0.00000\nimg_10.jpg,0.00000,0.00000,0.00000,0.99866,0.00004,0.00130,0.00000,0.00000,0.00000,0.00000</p>\n\n<p>2) Maybe the way scoring works is penalizing me somehow. I would image the sample submission of all 0.1 probabilities would be the worst type of submission. But, perhaps my strategy of just submitting the probabilities isn't the correct approach. </p>\n\n<p>3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. </p>\n\n<p>4) Maybe my model is garbage. But, I don't understand why my testing results show around 97% but the LB score is higher than the sample submission LB score. </p>\n\n<p>Thanks for helping!\nChris</p>",
  "messages": [
    {
      "id": 119035,
      "postDate": "2016-05-06T20:16:28.393Z",
      "content": "<p>Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">formula</a> of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.</p>\n\n<p>I recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.</p>",
      "rawMarkdown": "Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the [formula][1] of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.\r\n\r\nI recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation",
      "votes": 3
    },
    {
      "id": 119528,
      "postDate": "2016-05-11T07:18:10.510Z",
      "content": "<p>[quote=ckleban;119043]</p>\n\n<p>[quote=Jiao Dong;119035]</p>\n\n<p>Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">formula</a> of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.</p>\n\n<p>I recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?</p>\n\n<p>Also, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. </p>\n\n<p>[/quote]</p>\n\n<p>The problem is that currently you are giving <em>too little</em> weight to the other labels. It seems that you have overtrained a lot, meaning that your model is dead set on one label and when it makes a false prediction this means the score is penalised very heavily.</p>",
      "rawMarkdown": "[quote=ckleban;119043]\r\n\r\n[quote=Jiao Dong;119035]\r\n\r\nAccuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the [formula][1] of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.\r\n\r\nI recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\r\n\r\n[/quote]\r\n\r\nThanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?\r\n\r\nAlso, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. \r\n\r\n\r\n\r\n[/quote]\r\n\r\nThe problem is that currently you are giving *too little* weight to the other labels. It seems that you have overtrained a lot, meaning that your model is dead set on one label and when it makes a false prediction this means the score is penalised very heavily.\r\n",
      "votes": 1
    },
    {
      "id": 119031,
      "postDate": "2016-05-06T19:22:17.857Z",
      "content": "<p>You may try to split by drivers as discussed in <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras/114427#post114427\">here</a> instead of random splitting. </p>\n\n<p>[quote=ckleban;119025]</p>\n\n<p>3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. </p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "You may try to split by drivers as discussed in [here][1] instead of random splitting. \r\n\r\n\r\n[quote=ckleban;119025]\r\n\r\n\r\n3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. \r\n\r\n[/quote]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras/114427#post114427",
      "votes": 1
    },
    {
      "id": 119025,
      "postDate": "2016-05-06T18:46:42.897Z",
      "content": "<p>Hello,</p>\n\n<p>I'm new to Kaggle and ML and I'm running to an issue that I don't understand. </p>\n\n<p>When I submit the 'example submission' what has all 0.1 values, I get a score of about 2.3. When I train a simple DNN using tensorflow and skflow, I get a Test Accuracy of 0.97681 when comparing my trained model on my CV data. However, my LB score for this DNN is higher than the sample submission, at around 7 or higher. I'm not sure what could be causing this. My thoughts are:</p>\n\n<p>1) Maybe my submissions are wrong somehow? I'm currently submitting the probabilities, versus the predictions. Is that what I'm suppose to do? Fwiw, I'm using this function: classifier.predict_proba(X_test) to generate a 2D array of probabilities for the 70k+ test images. A small excerpt of my submission file looks like: </p>\n\n<p>img,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9\nimg_1.jpg,0.00000,0.00000,0.00034,0.00000,0.00000,0.99953,0.00013,0.00000,0.00000,0.00000\nimg_10.jpg,0.00000,0.00000,0.00000,0.99866,0.00004,0.00130,0.00000,0.00000,0.00000,0.00000</p>\n\n<p>2) Maybe the way scoring works is penalizing me somehow. I would image the sample submission of all 0.1 probabilities would be the worst type of submission. But, perhaps my strategy of just submitting the probabilities isn't the correct approach. </p>\n\n<p>3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. </p>\n\n<p>4) Maybe my model is garbage. But, I don't understand why my testing results show around 97% but the LB score is higher than the sample submission LB score. </p>\n\n<p>Thanks for helping!\nChris</p>",
      "rawMarkdown": "Hello,\r\n\r\nI'm new to Kaggle and ML and I'm running to an issue that I don't understand. \r\n\r\nWhen I submit the 'example submission' what has all 0.1 values, I get a score of about 2.3. When I train a simple DNN using tensorflow and skflow, I get a Test Accuracy of 0.97681 when comparing my trained model on my CV data. However, my LB score for this DNN is higher than the sample submission, at around 7 or higher. I'm not sure what could be causing this. My thoughts are:\r\n\r\n1) Maybe my submissions are wrong somehow? I'm currently submitting the probabilities, versus the predictions. Is that what I'm suppose to do? Fwiw, I'm using this function: classifier.predict_proba(X_test) to generate a 2D array of probabilities for the 70k+ test images. A small excerpt of my submission file looks like: \r\n\r\nimg,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9\r\nimg_1.jpg,0.00000,0.00000,0.00034,0.00000,0.00000,0.99953,0.00013,0.00000,0.00000,0.00000\r\nimg_10.jpg,0.00000,0.00000,0.00000,0.99866,0.00004,0.00130,0.00000,0.00000,0.00000,0.00000\r\n\r\n2) Maybe the way scoring works is penalizing me somehow. I would image the sample submission of all 0.1 probabilities would be the worst type of submission. But, perhaps my strategy of just submitting the probabilities isn't the correct approach. \r\n\r\n3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. \r\n\r\n4) Maybe my model is garbage. But, I don't understand why my testing results show around 97% but the LB score is higher than the sample submission LB score. \r\n\r\nThanks for helping!\r\nChris\r\n\r\n\r\n\r\n\r\n ",
      "votes": 1
    },
    {
      "id": 119531,
      "postDate": "2016-05-11T07:50:32.407Z",
      "content": "<p>[quote=ckleban;119042]</p>\n\n<p>[quote=the1owl;119032]</p>\n\n<p>Replace all your zeros with 0.1</p>\n\n<p>[/quote]</p>\n\n<p>Why would this help? If my model is picking the correct answer, why would I want to do this?</p>\n\n<p>[/quote]</p>\n\n<p>It would help because you are being very heavily penalised when your model gets it wrong by predicting 0 probability for the correct class, due to logloss metric.</p>\n\n<p>A simpler and perhaps more satisfying approach would be to increase the number of significant digits in your submission. You will get penalised much less for 0.000001 prediction than 0.000000, in cases where your model has got it wrong. I'd use maybe 9 decimal places.</p>\n\n<p>It does look like you have overfit to the training set though. You really need to split by driver to get meaningful CV numbers - splitting randomly by example will drive you to overfit, and the CV score you get will be unrelated to the LB score.</p>",
      "rawMarkdown": "[quote=ckleban;119042]\r\n\r\n[quote=the1owl;119032]\r\n\r\nReplace all your zeros with 0.1\r\n\r\n[/quote]\r\n\r\nWhy would this help? If my model is picking the correct answer, why would I want to do this?\r\n\r\n[/quote]\r\n\r\nIt would help because you are being very heavily penalised when your model gets it wrong by predicting 0 probability for the correct class, due to logloss metric.\r\n\r\nA simpler and perhaps more satisfying approach would be to increase the number of significant digits in your submission. You will get penalised much less for 0.000001 prediction than 0.000000, in cases where your model has got it wrong. I'd use maybe 9 decimal places.\r\n\r\nIt does look like you have overfit to the training set though. You really need to split by driver to get meaningful CV numbers - splitting randomly by example will drive you to overfit, and the CV score you get will be unrelated to the LB score.\r\n\r\n\r\n",
      "votes": 2
    },
    {
      "id": 122742,
      "postDate": "2016-06-06T22:06:37.817Z",
      "content": "<p>There are very few drivers and ~80 images per every class per every driver. So if the split is random, you will definitely have images from the same class and the same driver both in training and validation sets which are very similar. </p>\n\n<p>Check the image here: <a href=\"https://www.kaggle.com/hrantkhachatrian/state-farm-distracted-driver-detection/exploratory-classes-per-driver\">https://www.kaggle.com/hrantkhachatrian/state-farm-distracted-driver-detection/exploratory-classes-per-driver</a></p>\n\n<p>Basically, the training set is already &quot;augmented&quot; :) </p>",
      "rawMarkdown": "There are very few drivers and ~80 images per every class per every driver. So if the split is random, you will definitely have images from the same class and the same driver both in training and validation sets which are very similar. \r\n\r\nCheck the image here: https://www.kaggle.com/hrantkhachatrian/state-farm-distracted-driver-detection/exploratory-classes-per-driver\r\n\r\nBasically, the training set is already \"augmented\" :) "
    },
    {
      "id": 121051,
      "postDate": "2016-05-23T09:18:51.677Z",
      "content": "<p>[quote=ckleban;119042]</p>\n\n<p>[quote=the1owl;119032]</p>\n\n<p>Replace all your zeros with 0.1</p>\n\n<p>[/quote]</p>\n\n<p>Why would this help? If my model is picking the correct answer, why would I want to do this?</p>\n\n<p>[/quote]</p>\n\n<p>Because LogLoss tends to infinite if probability is in {0,1} and is wrong</p>",
      "rawMarkdown": "[quote=ckleban;119042]\r\n\r\n[quote=the1owl;119032]\r\n\r\nReplace all your zeros with 0.1\r\n\r\n[/quote]\r\n\r\nWhy would this help? If my model is picking the correct answer, why would I want to do this?\r\n\r\n[/quote]\r\n\r\n\r\nBecause LogLoss tends to infinite if probability is in {0,1} and is wrong"
    },
    {
      "id": 120288,
      "postDate": "2016-05-17T04:14:01.100Z",
      "content": "<p>Thanks everyone for your pointers. </p>\n\n<p>After splitting the training and CV data by driver, I indeed was overfitting in a big way. I've been using tensorflows's DNN and Convnets and both had overfitting issues. I'll be trying out hyper parameter tuning as well as try on pre-trained models (ie transfer learning) as my next step to see if I can get a model that works well with new, unseen drivers. </p>\n\n<p>--Chris</p>",
      "rawMarkdown": "Thanks everyone for your pointers. \r\n\r\nAfter splitting the training and CV data by driver, I indeed was overfitting in a big way. I've been using tensorflows's DNN and Convnets and both had overfitting issues. I'll be trying out hyper parameter tuning as well as try on pre-trained models (ie transfer learning) as my next step to see if I can get a model that works well with new, unseen drivers. \r\n\r\n--Chris\r\n"
    },
    {
      "id": 120028,
      "postDate": "2016-05-14T18:23:53.007Z",
      "content": "<p>[quote=paulgamble;120023]</p>\n\n<p>Neil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? </p>\n\n<p>[/quote]</p>\n\n<p>Yes. </p>\n\n<p>Generalisation of the classifier to new drivers is the goal of the training as tested by the LB.</p>\n\n<p>Optimising CV scores against a simple random split will significantly over-estimate performance, and it would be very hard to tell whether you were over-fitting for that goal.</p>",
      "rawMarkdown": "[quote=paulgamble;120023]\r\n\r\nNeil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? \r\n\r\n[/quote]\r\n\r\nYes. \r\n\r\nGeneralisation of the classifier to new drivers is the goal of the training as tested by the LB.\r\n\r\nOptimising CV scores against a simple random split will significantly over-estimate performance, and it would be very hard to tell whether you were over-fitting for that goal.\r\n\r\n"
    },
    {
      "id": 120023,
      "postDate": "2016-05-14T17:49:37.733Z",
      "content": "<p>Neil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? </p>",
      "rawMarkdown": "Neil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? "
    },
    {
      "id": 119043,
      "postDate": "2016-05-06T22:05:55.810Z",
      "content": "<p>[quote=Jiao Dong;119035]</p>\n\n<p>Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">formula</a> of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.</p>\n\n<p>I recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?</p>\n\n<p>Also, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. </p>",
      "rawMarkdown": "[quote=Jiao Dong;119035]\r\n\r\nAccuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the [formula][1] of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.\r\n\r\nI recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\r\n\r\n[/quote]\r\n\r\nThanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?\r\n\r\nAlso, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. \r\n\r\n"
    },
    {
      "id": 119042,
      "postDate": "2016-05-06T22:02:29.403Z",
      "content": "<p>[quote=the1owl;119032]</p>\n\n<p>Replace all your zeros with 0.1</p>\n\n<p>[/quote]</p>\n\n<p>Why would this help? If my model is picking the correct answer, why would I want to do this?</p>",
      "rawMarkdown": "[quote=the1owl;119032]\r\n\r\nReplace all your zeros with 0.1\r\n\r\n[/quote]\r\n\r\nWhy would this help? If my model is picking the correct answer, why would I want to do this?"
    },
    {
      "id": 119032,
      "postDate": "2016-05-06T19:37:29.753Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 119035,
      "author_name": "Jiao Dong",
      "author_url": "",
      "post_date": "2016-05-06T20:16:28.393000",
      "content": "<p>Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">formula</a> of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.</p>\n\n<p>I recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 119528,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2016-05-11T07:18:10.510000",
      "content": "<p>[quote=ckleban;119043]</p>\n\n<p>[quote=Jiao Dong;119035]</p>\n\n<p>Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">formula</a> of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.</p>\n\n<p>I recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?</p>\n\n<p>Also, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. </p>\n\n<p>[/quote]</p>\n\n<p>The problem is that currently you are giving <em>too little</em> weight to the other labels. It seems that you have overtrained a lot, meaning that your model is dead set on one label and when it makes a false prediction this means the score is penalised very heavily.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119031,
      "author_name": "june",
      "author_url": "",
      "post_date": "2016-05-06T19:22:17.857000",
      "content": "<p>You may try to split by drivers as discussed in <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras/114427#post114427\">here</a> instead of random splitting. </p>\n\n<p>[quote=ckleban;119025]</p>\n\n<p>3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. </p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 119531,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-05-11T07:50:32.407000",
      "content": "<p>[quote=ckleban;119042]</p>\n\n<p>[quote=the1owl;119032]</p>\n\n<p>Replace all your zeros with 0.1</p>\n\n<p>[/quote]</p>\n\n<p>Why would this help? If my model is picking the correct answer, why would I want to do this?</p>\n\n<p>[/quote]</p>\n\n<p>It would help because you are being very heavily penalised when your model gets it wrong by predicting 0 probability for the correct class, due to logloss metric.</p>\n\n<p>A simpler and perhaps more satisfying approach would be to increase the number of significant digits in your submission. You will get penalised much less for 0.000001 prediction than 0.000000, in cases where your model has got it wrong. I'd use maybe 9 decimal places.</p>\n\n<p>It does look like you have overfit to the training set though. You really need to split by driver to get meaningful CV numbers - splitting randomly by example will drive you to overfit, and the CV score you get will be unrelated to the LB score.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 122742,
      "author_name": "Hrant Khachatrian",
      "author_url": "",
      "post_date": "2016-06-06T22:06:37.817000",
      "content": "<p>There are very few drivers and ~80 images per every class per every driver. So if the split is random, you will definitely have images from the same class and the same driver both in training and validation sets which are very similar. </p>\n\n<p>Check the image here: <a href=\"https://www.kaggle.com/hrantkhachatrian/state-farm-distracted-driver-detection/exploratory-classes-per-driver\">https://www.kaggle.com/hrantkhachatrian/state-farm-distracted-driver-detection/exploratory-classes-per-driver</a></p>\n\n<p>Basically, the training set is already &quot;augmented&quot; :) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 121051,
      "author_name": "Shahnawaz Akhtar",
      "author_url": "",
      "post_date": "2016-05-23T09:18:51.677000",
      "content": "<p>[quote=ckleban;119042]</p>\n\n<p>[quote=the1owl;119032]</p>\n\n<p>Replace all your zeros with 0.1</p>\n\n<p>[/quote]</p>\n\n<p>Why would this help? If my model is picking the correct answer, why would I want to do this?</p>\n\n<p>[/quote]</p>\n\n<p>Because LogLoss tends to infinite if probability is in {0,1} and is wrong</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120288,
      "author_name": "ckleban",
      "author_url": "",
      "post_date": "2016-05-17T04:14:01.100000",
      "content": "<p>Thanks everyone for your pointers. </p>\n\n<p>After splitting the training and CV data by driver, I indeed was overfitting in a big way. I've been using tensorflows's DNN and Convnets and both had overfitting issues. I'll be trying out hyper parameter tuning as well as try on pre-trained models (ie transfer learning) as my next step to see if I can get a model that works well with new, unseen drivers. </p>\n\n<p>--Chris</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120028,
      "author_name": "Neil Slater",
      "author_url": "",
      "post_date": "2016-05-14T18:23:53.007000",
      "content": "<p>[quote=paulgamble;120023]</p>\n\n<p>Neil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? </p>\n\n<p>[/quote]</p>\n\n<p>Yes. </p>\n\n<p>Generalisation of the classifier to new drivers is the goal of the training as tested by the LB.</p>\n\n<p>Optimising CV scores against a simple random split will significantly over-estimate performance, and it would be very hard to tell whether you were over-fitting for that goal.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 120023,
      "author_name": "gowdy",
      "author_url": "",
      "post_date": "2016-05-14T17:49:37.733000",
      "content": "<p>Neil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119043,
      "author_name": "ckleban",
      "author_url": "",
      "post_date": "2016-05-06T22:05:55.810000",
      "content": "<p>[quote=Jiao Dong;119035]</p>\n\n<p>Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\">formula</a> of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.</p>\n\n<p>I recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.</p>\n\n<p>[/quote]</p>\n\n<p>Thanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?</p>\n\n<p>Also, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119042,
      "author_name": "ckleban",
      "author_url": "",
      "post_date": "2016-05-06T22:02:29.403000",
      "content": "<p>[quote=the1owl;119032]</p>\n\n<p>Replace all your zeros with 0.1</p>\n\n<p>[/quote]</p>\n\n<p>Why would this help? If my model is picking the correct answer, why would I want to do this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 119032,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-05-06T19:37:29.753000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119035": "Accuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the [formula][1] of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.\r\n\r\nI recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation",
    "119528": "[quote=ckleban;119043]\r\n\r\n[quote=Jiao Dong;119035]\r\n\r\nAccuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the [formula][1] of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.\r\n\r\nI recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\r\n\r\n[/quote]\r\n\r\nThanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?\r\n\r\nAlso, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. \r\n\r\n\r\n\r\n[/quote]\r\n\r\nThe problem is that currently you are giving *too little* weight to the other labels. It seems that you have overtrained a lot, meaning that your model is dead set on one label and when it makes a false prediction this means the score is penalised very heavily.\r\n",
    "119031": "You may try to split by drivers as discussed in [here][1] instead of random splitting. \r\n\r\n\r\n[quote=ckleban;119025]\r\n\r\n\r\n3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. \r\n\r\n[/quote]\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/19971/simple-solution-keras/114427#post114427",
    "119025": "Hello,\r\n\r\nI'm new to Kaggle and ML and I'm running to an issue that I don't understand. \r\n\r\nWhen I submit the 'example submission' what has all 0.1 values, I get a score of about 2.3. When I train a simple DNN using tensorflow and skflow, I get a Test Accuracy of 0.97681 when comparing my trained model on my CV data. However, my LB score for this DNN is higher than the sample submission, at around 7 or higher. I'm not sure what could be causing this. My thoughts are:\r\n\r\n1) Maybe my submissions are wrong somehow? I'm currently submitting the probabilities, versus the predictions. Is that what I'm suppose to do? Fwiw, I'm using this function: classifier.predict_proba(X_test) to generate a 2D array of probabilities for the 70k+ test images. A small excerpt of my submission file looks like: \r\n\r\nimg,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9\r\nimg_1.jpg,0.00000,0.00000,0.00034,0.00000,0.00000,0.99953,0.00013,0.00000,0.00000,0.00000\r\nimg_10.jpg,0.00000,0.00000,0.00000,0.99866,0.00004,0.00130,0.00000,0.00000,0.00000,0.00000\r\n\r\n2) Maybe the way scoring works is penalizing me somehow. I would image the sample submission of all 0.1 probabilities would be the worst type of submission. But, perhaps my strategy of just submitting the probabilities isn't the correct approach. \r\n\r\n3) Perhaps I'm splitting my training data incorrectly into train and CV data sets? I'm currently splitting the training data randomly into a train (90%) and CV data set (10%), but I'm not doing things like ensuring the CV data is of drivers not seen in the training data. My thought is that while I might be able to improve things here, I wouldn't imagine getting a worse score than the sample submission. \r\n\r\n4) Maybe my model is garbage. But, I don't understand why my testing results show around 97% but the LB score is higher than the sample submission LB score. \r\n\r\nThanks for helping!\r\nChris\r\n\r\n\r\n\r\n\r\n ",
    "119531": "[quote=ckleban;119042]\r\n\r\n[quote=the1owl;119032]\r\n\r\nReplace all your zeros with 0.1\r\n\r\n[/quote]\r\n\r\nWhy would this help? If my model is picking the correct answer, why would I want to do this?\r\n\r\n[/quote]\r\n\r\nIt would help because you are being very heavily penalised when your model gets it wrong by predicting 0 probability for the correct class, due to logloss metric.\r\n\r\nA simpler and perhaps more satisfying approach would be to increase the number of significant digits in your submission. You will get penalised much less for 0.000001 prediction than 0.000000, in cases where your model has got it wrong. I'd use maybe 9 decimal places.\r\n\r\nIt does look like you have overfit to the training set though. You really need to split by driver to get meaningful CV numbers - splitting randomly by example will drive you to overfit, and the CV score you get will be unrelated to the LB score.\r\n\r\n\r\n",
    "122742": "There are very few drivers and ~80 images per every class per every driver. So if the split is random, you will definitely have images from the same class and the same driver both in training and validation sets which are very similar. \r\n\r\nCheck the image here: https://www.kaggle.com/hrantkhachatrian/state-farm-distracted-driver-detection/exploratory-classes-per-driver\r\n\r\nBasically, the training set is already \"augmented\" :) ",
    "121051": "[quote=ckleban;119042]\r\n\r\n[quote=the1owl;119032]\r\n\r\nReplace all your zeros with 0.1\r\n\r\n[/quote]\r\n\r\nWhy would this help? If my model is picking the correct answer, why would I want to do this?\r\n\r\n[/quote]\r\n\r\n\r\nBecause LogLoss tends to infinite if probability is in {0,1} and is wrong",
    "120288": "Thanks everyone for your pointers. \r\n\r\nAfter splitting the training and CV data by driver, I indeed was overfitting in a big way. I've been using tensorflows's DNN and Convnets and both had overfitting issues. I'll be trying out hyper parameter tuning as well as try on pre-trained models (ie transfer learning) as my next step to see if I can get a model that works well with new, unseen drivers. \r\n\r\n--Chris\r\n",
    "120028": "[quote=paulgamble;120023]\r\n\r\nNeil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? \r\n\r\n[/quote]\r\n\r\nYes. \r\n\r\nGeneralisation of the classifier to new drivers is the goal of the training as tested by the LB.\r\n\r\nOptimising CV scores against a simple random split will significantly over-estimate performance, and it would be very hard to tell whether you were over-fitting for that goal.\r\n\r\n",
    "120023": "Neil, just to clarify: you're recommending that CV folds consist of entire drivers - rather than a random subset of training images. Is this because the test set has entirely different drivers from the training set? ",
    "119043": "[quote=Jiao Dong;119035]\r\n\r\nAccuracy is just one of the indicator of your model, in fact you can design many training strategies that would eventually converge to high training accuracy. But your submission is evaluated based on loss. According to the [formula][1] of evaluation, in order to get smaller testing loss you would like to assign bigger probability to the correct label. However there will be some cases you made the wrong prediction,  since you aggressively set most probabilities to 0, after rescaling weights of your submission, you would take a bigger penalty in loss as well.\r\n\r\nI recommend to focus on your training loss, especially your cross-validation loss during training, they would be better metrics to give you hint about final submission loss score.\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/state-farm-distracted-driver-detection/details/evaluation\r\n\r\n[/quote]\r\n\r\nThanks. My cross validation loss is something like 0.127. I assume I should be getting this number a lot smaller?\r\n\r\nAlso, should I refactor my submissions to give more weight to my predicted results and perhaps less weight to the other labels? Or should I focus on my CV loss. \r\n\r\n",
    "119042": "[quote=the1owl;119032]\r\n\r\nReplace all your zeros with 0.1\r\n\r\n[/quote]\r\n\r\nWhy would this help? If my model is picking the correct answer, why would I want to do this?",
    "119032": ""
  }
}