{
  "id": 12740,
  "title": "CV scores vs Leaderboard",
  "url": "/competitions/malware-classification/discussion/12740",
  "author_name": "",
  "post_date": "2015-03-09T10:17:42.853Z",
  "votes": null,
  "comment_count": 15,
  "views": 3361,
  "content": "<p>How does&nbsp;your local CV score compare against the public LB?<br><br>I'm asking because my CV scores are&nbsp;directionally correct but seem to be consistently higher (around +0.005) than the LB.</p>",
  "messages": [
    {
      "id": "65767",
      "postDate": "03/09/2015 10:17:42",
      "content": "<p>How does&nbsp;your local CV score compare against the public LB?<br><br>I'm asking because my CV scores are&nbsp;directionally correct but seem to be consistently higher (around +0.005) than the LB.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65778",
      "postDate": "03/09/2015 14:14:39",
      "content": "<p>my 4-fold CV is about 0.0007~0.001 higher&nbsp;than LB</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65810",
      "postDate": "03/09/2015 19:32:39",
      "content": "<p>@Stergios,&nbsp;</p>\n<p>Mine (10-fold) CV score compared to LB is similar to your situation.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65815",
      "postDate": "03/09/2015 21:28:04",
      "content": "<p>I'm getting&nbsp;0.02315174 in a 5-fold cv, but my LB score is ~0.109...i'm not sure why =p</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65817",
      "postDate": "03/09/2015 22:37:15",
      "content": "<p>Mine is also about 0.001 -- 0.002 better in CV results.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65832",
      "postDate": "03/10/2015 01:27:17",
      "content": "<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65839",
      "postDate": "03/10/2015 02:26:11",
      "content": "<p>[quote=Jiming Ye;65832]</p>\n<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>\n<p>[/quote]</p>\n<p>Hi, I use the code below, where y_true is Nx1 true label vector with value ranging from 0 to 8, and y_pred is Nx9 prediction arrays.</p>\n<p>import numpy as np</p>\n<p>from math import log, exp, sqrt,factorial<br>def multiclass_log_loss(y_true, y_pred, eps=1e-15):</p>\n<p>&nbsp; &nbsp; &nbsp; predictions = np.clip(y_pred, eps, 1 - eps)</p>\n<p>&nbsp; &nbsp; &nbsp; # normalize row sums to 1<br>&nbsp; &nbsp; &nbsp; predictions /= predictions.sum(axis=1)[:, np.newaxis]</p>\n<p>&nbsp; &nbsp; &nbsp; actual = np.zeros(y_pred.shape)<br>&nbsp; &nbsp; &nbsp; n_samples = actual.shape[0]<br><br>&nbsp; &nbsp; &nbsp; actual[np.arange(n_samples), y_true.astype(int)] = 1<br>&nbsp; &nbsp; &nbsp; vectsum = np.sum(actual * np.log(predictions))<br>&nbsp; &nbsp; &nbsp; loss = -1.0 / n_samples * vectsum<br>&nbsp; &nbsp; &nbsp; return loss</p>\n<p>I steal it from the tutorial of the Science Bowl contest.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65840",
      "postDate": "03/10/2015 02:28:34",
      "content": "<p>Not sure if this is what you're asking... In Python:</p>\n<p>from sklearn.metrics import log_loss</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65841",
      "postDate": "03/10/2015 02:28:51",
      "content": "<p>[quote=Bats &amp; Robots;65815]</p>\n<p>I'm getting&nbsp;0.02315174 in a 5-fold cv, but my LB score is ~0.109...i'm not sure why =p</p>\n<p>[/quote]</p>\n<p>you may want to check your test data. they can be flawed during downloading.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65842",
      "postDate": "03/10/2015 02:40:08",
      "content": "<p>[quote=rcarson;65839]</p>\n<p>[quote=Jiming Ye;65832]</p>\n<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>\n<p>[/quote]</p>\n<p>Hi, I use the code below, where y_true is Nx1 true label vector with value ranging from 0 to 8, and y_pred is Nx9 prediction arrays.</p>\n<p>import numpy as np</p>\n<p>from math import log, exp, sqrt,factorial<br>def multiclass_log_loss(y_true, y_pred, eps=1e-15):</p>\n<p>&nbsp; &nbsp; &nbsp; predictions = np.clip(y_pred, eps, 1 - eps)</p>\n<p>&nbsp; &nbsp; &nbsp; # normalize row sums to 1<br>&nbsp; &nbsp; &nbsp; predictions /= predictions.sum(axis=1)[:, np.newaxis]</p>\n<p>&nbsp; &nbsp; &nbsp; actual = np.zeros(y_pred.shape)<br>&nbsp; &nbsp; &nbsp; n_samples = actual.shape[0]<br><br>&nbsp; &nbsp; &nbsp; actual[np.arange(n_samples), y_true.astype(int)] = 1<br>&nbsp; &nbsp; &nbsp; vectsum = np.sum(actual * np.log(predictions))<br>&nbsp; &nbsp; &nbsp; loss = -1.0 / n_samples * vectsum<br>&nbsp; &nbsp; &nbsp; return loss</p>\n<p>I steal it from the tutorial of the Science Bowl contest.</p>\n<p>[/quote] I did exactly the same thing!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65845",
      "postDate": "03/10/2015 02:43:18",
      "content": "<p>[quote=Little Boat;65842]</p>\n<p>I did exactly the same thing!</p>\n<p>[/quote]</p>\n<p>Well done! gogogo :D</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65855",
      "postDate": "03/10/2015 03:50:20",
      "content": "<p>[quote=rcarson;65839]</p>\n<p>[quote=Jiming Ye;65832]</p>\n<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>\n<p>[/quote]</p>\n<p>Hi, I use the code below, where y_true is Nx1 true label vector with value ranging from 0 to 8, and y_pred is Nx9 prediction arrays.</p>\n<p>import numpy as np</p>\n<p>from math import log, exp, sqrt,factorial<br>def multiclass_log_loss(y_true, y_pred, eps=1e-15):</p>\n<p>&nbsp; &nbsp; &nbsp; predictions = np.clip(y_pred, eps, 1 - eps)</p>\n<p>&nbsp; &nbsp; &nbsp; # normalize row sums to 1<br>&nbsp; &nbsp; &nbsp; predictions /= predictions.sum(axis=1)[:, np.newaxis]</p>\n<p>&nbsp; &nbsp; &nbsp; actual = np.zeros(y_pred.shape)<br>&nbsp; &nbsp; &nbsp; n_samples = actual.shape[0]<br><br>&nbsp; &nbsp; &nbsp; actual[np.arange(n_samples), y_true.astype(int)] = 1<br>&nbsp; &nbsp; &nbsp; vectsum = np.sum(actual * np.log(predictions))<br>&nbsp; &nbsp; &nbsp; loss = -1.0 / n_samples * vectsum<br>&nbsp; &nbsp; &nbsp; return loss</p>\n<p>I steal it from the tutorial of the Science Bowl contest.</p>\n<p>[/quote]</p>\n<p>Great! Thanks so much</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65979",
      "postDate": "03/11/2015 20:28:46",
      "content": "<p>[quote=rcarson;65841]</p>\n<p>[quote=Bats &amp; Robots;65815]</p>\n<p>I'm getting&nbsp;0.02315174 in a 5-fold cv, but my LB score is ~0.109...i'm not sure why =p</p>\n<p>[/quote]</p>\n<p>you may want to check your test data. they can be flawed during downloading.</p>\n<p>[/quote]</p>\n\n<p>I did check the md5 hashes, they are ok. I added some features, then my new CV got&nbsp;0.0196 but now i got 0.11 as LB. I'm using SVM...any hints?</p>\n<p>thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65985",
      "postDate": "03/11/2015 21:13:37",
      "content": "<p>0.11 is actually low for any kind of messing up of data or id or label stuff I could think of. it seems more like overfitting to me now. Anyway maybe you could check whether the distribution of each class in submission is similar to the training data set. Or manually check if there are any weird thing in submission file. Ideally for a 0.0196 cv you should see that most files are predicted &gt;0.99 probability to only one class.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65988",
      "postDate": "03/11/2015 21:45:44",
      "content": "<p>[quote=rcarson;65985]</p>\n<p>0.11 is actually low for any kind of messing up of data or id or label stuff I could think of. it seems more like overfitting to me now. Anyway maybe you could check whether the distribution of each class in submission is similar to the training data set. Or manually check if there are any weird thing in submission file. Ideally for a 0.0196 cv you should see that most files are predicted &gt;0.99 probability to only one class.&nbsp;</p>\n<p>[/quote]</p>\n<p>I&nbsp;managed to get ~0.08 on LB after removing some weird normalization and&nbsp;changing models. Now i'll re-run the SVM without this normalization mess (since it will perform scaling by default). Also the submission file has a similar distribution as the train set. Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "66019",
      "postDate": "03/12/2015 08:16:40",
      "content": "<p>[quote=Bats &amp; Robots;65979]</p>\n<p>I did check the md5 hashes, they are ok. I added some features, then my new CV got&nbsp;0.0196 but now i got 0.11 as LB. I'm using SVM...any hints?</p>\n<p>thanks!</p>\n<p>[/quote]</p>\n\n<p>If it's not overfitting, it could be some mistake in the CV process. Take care to not use the validation data during training in any way.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 65778,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/09/2015 14:14:39",
      "content": "<p>my 4-fold CV is about 0.0007~0.001 higher&nbsp;than LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65810,
      "author_name": "jieqchen",
      "author_url": "",
      "post_date": "03/09/2015 19:32:39",
      "content": "<p>@Stergios,&nbsp;</p>\n<p>Mine (10-fold) CV score compared to LB is similar to your situation.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65815,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "03/09/2015 21:28:04",
      "content": "<p>I'm getting&nbsp;0.02315174 in a 5-fold cv, but my LB score is ~0.109...i'm not sure why =p</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65817,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "03/09/2015 22:37:15",
      "content": "<p>Mine is also about 0.001 -- 0.002 better in CV results.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65832,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "03/10/2015 01:27:17",
      "content": "<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65839,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/10/2015 02:26:11",
      "content": "<p>[quote=Jiming Ye;65832]</p>\n<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>\n<p>[/quote]</p>\n<p>Hi, I use the code below, where y_true is Nx1 true label vector with value ranging from 0 to 8, and y_pred is Nx9 prediction arrays.</p>\n<p>import numpy as np</p>\n<p>from math import log, exp, sqrt,factorial<br>def multiclass_log_loss(y_true, y_pred, eps=1e-15):</p>\n<p>&nbsp; &nbsp; &nbsp; predictions = np.clip(y_pred, eps, 1 - eps)</p>\n<p>&nbsp; &nbsp; &nbsp; # normalize row sums to 1<br>&nbsp; &nbsp; &nbsp; predictions /= predictions.sum(axis=1)[:, np.newaxis]</p>\n<p>&nbsp; &nbsp; &nbsp; actual = np.zeros(y_pred.shape)<br>&nbsp; &nbsp; &nbsp; n_samples = actual.shape[0]<br><br>&nbsp; &nbsp; &nbsp; actual[np.arange(n_samples), y_true.astype(int)] = 1<br>&nbsp; &nbsp; &nbsp; vectsum = np.sum(actual * np.log(predictions))<br>&nbsp; &nbsp; &nbsp; loss = -1.0 / n_samples * vectsum<br>&nbsp; &nbsp; &nbsp; return loss</p>\n<p>I steal it from the tutorial of the Science Bowl contest.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65840,
      "author_name": "jieqchen",
      "author_url": "",
      "post_date": "03/10/2015 02:28:34",
      "content": "<p>Not sure if this is what you're asking... In Python:</p>\n<p>from sklearn.metrics import log_loss</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65841,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/10/2015 02:28:51",
      "content": "<p>[quote=Bats &amp; Robots;65815]</p>\n<p>I'm getting&nbsp;0.02315174 in a 5-fold cv, but my LB score is ~0.109...i'm not sure why =p</p>\n<p>[/quote]</p>\n<p>you may want to check your test data. they can be flawed during downloading.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65842,
      "author_name": "xiaozhouwang",
      "author_url": "",
      "post_date": "03/10/2015 02:40:08",
      "content": "<p>[quote=rcarson;65839]</p>\n<p>[quote=Jiming Ye;65832]</p>\n<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>\n<p>[/quote]</p>\n<p>Hi, I use the code below, where y_true is Nx1 true label vector with value ranging from 0 to 8, and y_pred is Nx9 prediction arrays.</p>\n<p>import numpy as np</p>\n<p>from math import log, exp, sqrt,factorial<br>def multiclass_log_loss(y_true, y_pred, eps=1e-15):</p>\n<p>&nbsp; &nbsp; &nbsp; predictions = np.clip(y_pred, eps, 1 - eps)</p>\n<p>&nbsp; &nbsp; &nbsp; # normalize row sums to 1<br>&nbsp; &nbsp; &nbsp; predictions /= predictions.sum(axis=1)[:, np.newaxis]</p>\n<p>&nbsp; &nbsp; &nbsp; actual = np.zeros(y_pred.shape)<br>&nbsp; &nbsp; &nbsp; n_samples = actual.shape[0]<br><br>&nbsp; &nbsp; &nbsp; actual[np.arange(n_samples), y_true.astype(int)] = 1<br>&nbsp; &nbsp; &nbsp; vectsum = np.sum(actual * np.log(predictions))<br>&nbsp; &nbsp; &nbsp; loss = -1.0 / n_samples * vectsum<br>&nbsp; &nbsp; &nbsp; return loss</p>\n<p>I steal it from the tutorial of the Science Bowl contest.</p>\n<p>[/quote] I did exactly the same thing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65845,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/10/2015 02:43:18",
      "content": "<p>[quote=Little Boat;65842]</p>\n<p>I did exactly the same thing!</p>\n<p>[/quote]</p>\n<p>Well done! gogogo :D</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65855,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "03/10/2015 03:50:20",
      "content": "<p>[quote=rcarson;65839]</p>\n<p>[quote=Jiming Ye;65832]</p>\n<p>How do you guys get your CV LogLoss? I just got accuracies. About 99%</p>\n<p>[/quote]</p>\n<p>Hi, I use the code below, where y_true is Nx1 true label vector with value ranging from 0 to 8, and y_pred is Nx9 prediction arrays.</p>\n<p>import numpy as np</p>\n<p>from math import log, exp, sqrt,factorial<br>def multiclass_log_loss(y_true, y_pred, eps=1e-15):</p>\n<p>&nbsp; &nbsp; &nbsp; predictions = np.clip(y_pred, eps, 1 - eps)</p>\n<p>&nbsp; &nbsp; &nbsp; # normalize row sums to 1<br>&nbsp; &nbsp; &nbsp; predictions /= predictions.sum(axis=1)[:, np.newaxis]</p>\n<p>&nbsp; &nbsp; &nbsp; actual = np.zeros(y_pred.shape)<br>&nbsp; &nbsp; &nbsp; n_samples = actual.shape[0]<br><br>&nbsp; &nbsp; &nbsp; actual[np.arange(n_samples), y_true.astype(int)] = 1<br>&nbsp; &nbsp; &nbsp; vectsum = np.sum(actual * np.log(predictions))<br>&nbsp; &nbsp; &nbsp; loss = -1.0 / n_samples * vectsum<br>&nbsp; &nbsp; &nbsp; return loss</p>\n<p>I steal it from the tutorial of the Science Bowl contest.</p>\n<p>[/quote]</p>\n<p>Great! Thanks so much</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65979,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "03/11/2015 20:28:46",
      "content": "<p>[quote=rcarson;65841]</p>\n<p>[quote=Bats &amp; Robots;65815]</p>\n<p>I'm getting&nbsp;0.02315174 in a 5-fold cv, but my LB score is ~0.109...i'm not sure why =p</p>\n<p>[/quote]</p>\n<p>you may want to check your test data. they can be flawed during downloading.</p>\n<p>[/quote]</p>\n\n<p>I did check the md5 hashes, they are ok. I added some features, then my new CV got&nbsp;0.0196 but now i got 0.11 as LB. I'm using SVM...any hints?</p>\n<p>thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65985,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/11/2015 21:13:37",
      "content": "<p>0.11 is actually low for any kind of messing up of data or id or label stuff I could think of. it seems more like overfitting to me now. Anyway maybe you could check whether the distribution of each class in submission is similar to the training data set. Or manually check if there are any weird thing in submission file. Ideally for a 0.0196 cv you should see that most files are predicted &gt;0.99 probability to only one class.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65988,
      "author_name": "snowdog",
      "author_url": "",
      "post_date": "03/11/2015 21:45:44",
      "content": "<p>[quote=rcarson;65985]</p>\n<p>0.11 is actually low for any kind of messing up of data or id or label stuff I could think of. it seems more like overfitting to me now. Anyway maybe you could check whether the distribution of each class in submission is similar to the training data set. Or manually check if there are any weird thing in submission file. Ideally for a 0.0196 cv you should see that most files are predicted &gt;0.99 probability to only one class.&nbsp;</p>\n<p>[/quote]</p>\n<p>I&nbsp;managed to get ~0.08 on LB after removing some weird normalization and&nbsp;changing models. Now i'll re-run the SVM without this normalization mess (since it will perform scaling by default). Also the submission file has a similar distribution as the train set. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 66019,
      "author_name": "asterios",
      "author_url": "",
      "post_date": "03/12/2015 08:16:40",
      "content": "<p>[quote=Bats &amp; Robots;65979]</p>\n<p>I did check the md5 hashes, they are ok. I added some features, then my new CV got&nbsp;0.0196 but now i got 0.11 as LB. I'm using SVM...any hints?</p>\n<p>thanks!</p>\n<p>[/quote]</p>\n\n<p>If it's not overfitting, it could be some mistake in the CV process. Take care to not use the validation data during training in any way.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "65767": "",
    "65778": "",
    "65810": "",
    "65815": "",
    "65817": "",
    "65832": "",
    "65839": "",
    "65840": "",
    "65841": "",
    "65842": "",
    "65845": "",
    "65855": "",
    "65979": "",
    "65985": "",
    "65988": "",
    "66019": ""
  },
  "source": "meta"
}