{
  "id": 32400,
  "title": "Local validation vs. leaderboard: 12 vs 24",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/32400",
  "author_name": "",
  "post_date": "2017-05-02T07:06:41.559546600Z",
  "votes": 5,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Has anyone built a reliable local cross-validation strategy? For me the gap is massive, up to 12 on local vs 24 on LB.</p>\n\n<p>I'm leaving about 250 random images for the model that predicts counts (so the CNN does not see them at all). Then I'm doing a 5-fold cross-validation on this 250 images using a couple of models such as XGBoost and Lasso.</p>\n\n<p>What could explain such a huge gap? It could obviously be a bug in my code, or very different images in test vs. train (I didn't check predictions on test but the images look similar). Apart from that, I don't see any obvious reasons, or ways I'm leaking train into my validation.</p>",
  "messages": [
    {
      "id": "179595",
      "postDate": "05/02/2017 07:06:41",
      "content": "<p>Has anyone built a reliable local cross-validation strategy? For me the gap is massive, up to 12 on local vs 24 on LB.</p>\n\n<p>I'm leaving about 250 random images for the model that predicts counts (so the CNN does not see them at all). Then I'm doing a 5-fold cross-validation on this 250 images using a couple of models such as XGBoost and Lasso.</p>\n\n<p>What could explain such a huge gap? It could obviously be a bug in my code, or very different images in test vs. train (I didn't check predictions on test but the images look similar). Apart from that, I don't see any obvious reasons, or ways I'm leaking train into my validation.</p>",
      "rawMarkdown": "Has anyone built a reliable local cross-validation strategy? For me the gap is massive, up to 12 on local vs 24 on LB.\n\nI'm leaving about 250 random images for the model that predicts counts (so the CNN does not see them at all). Then I'm doing a 5-fold cross-validation on this 250 images using a couple of models such as XGBoost and Lasso.\n\nWhat could explain such a huge gap? It could obviously be a bug in my code, or very different images in test vs. train (I didn't check predictions on test but the images look similar). Apart from that, I don't see any obvious reasons, or ways I'm leaking train into my validation.",
      "votes": null
    },
    {
      "id": "179651",
      "postDate": "05/02/2017 12:15:08",
      "content": "<p>Konstantin, do you train a CNN on first ~750 images and than tune classifiers on hold-out 250 ?</p>",
      "rawMarkdown": "Konstantin, do you train a CNN on first ~750 images and than tune classifiers on hold-out 250 ?",
      "votes": null
    },
    {
      "id": "179657",
      "postDate": "05/02/2017 12:27:35",
      "content": "<p>Yes - actually, I'm training a CNN on about 650 images, and use about 100 for validation of the CNN, and then tune regression models that take CNN predictions as input on the hold-out 250.</p>\n\n<p>Edit: and I'm also excluding images from MismatchedTrainImages.txt + 941,  200, 491, 912</p>",
      "rawMarkdown": "Yes - actually, I'm training a CNN on about 650 images, and use about 100 for validation of the CNN, and then tune regression models that take CNN predictions as input on the hold-out 250.\n\nEdit: and I'm also excluding images from MismatchedTrainImages.txt + 941,  200, 491, 912",
      "votes": null
    },
    {
      "id": "179688",
      "postDate": "05/02/2017 14:11:23",
      "content": "<p>My current approach make massive gap too.<br>\nIn my opinion, we should not believe public LB.</p>",
      "rawMarkdown": "My current approach make massive gap too.<br>\nIn my opinion, we should not believe public LB.",
      "votes": null
    },
    {
      "id": "179767",
      "postDate": "05/02/2017 19:49:49",
      "content": "<blockquote>\n  <p>In my opinion, we should not believe public LB.</p>\n</blockquote>\n\n<p>Do you think that the private LB will be better? Do you have any hypothesis about that gap between local CV and public LB?</p>",
      "rawMarkdown": "> In my opinion, we should not believe public LB.\n\nDo you think that the private LB will be better? Do you have any hypothesis about that gap between local CV and public LB?",
      "votes": null
    },
    {
      "id": "179906",
      "postDate": "05/03/2017 08:14:47",
      "content": "<p>Some competitors point out that number of sea-lions in <code>train.csv</code> don't consistent real number. This implies <code>train.csv</code>　includes much noise. I don't know that private LB will be better, but I hope our final submission evaluated with real number.<br></p>\n\n<p>By the way, admin shared <code>MismatchedTrainImages.txt</code> after starting competition. I wonder if there is <code>MismatchedTestImages.txt</code> potentially or not.</p>",
      "rawMarkdown": "Some competitors point out that number of sea-lions in `train.csv` don't consistent real number. This implies `train.csv`　includes much noise. I don't know that private LB will be better, but I hope our final submission evaluated with real number.<br>\n\nBy the way, admin shared `MismatchedTrainImages.txt` after starting competition. I wonder if there is `MismatchedTestImages.txt` potentially or not.",
      "votes": null
    },
    {
      "id": "182904",
      "postDate": "05/16/2017 09:00:58",
      "content": "<p>Hello.\nHave the same problem =(\nI tried to submit all zeros, and it gave me the score 29.</p>",
      "rawMarkdown": "Hello.\nHave the same problem =(\nI tried to submit all zeros, and it gave me the score 29.",
      "votes": null
    },
    {
      "id": "182912",
      "postDate": "05/16/2017 09:44:20",
      "content": "<p>Just to understand your approach, and I apologize in advance if I misunderstood something; couldn't you also validate your model on the test set? That will only allow you to validate the counts, not the predicted coordinates of the animals, but wouldn't that be enough?</p>",
      "rawMarkdown": "Just to understand your approach, and I apologize in advance if I misunderstood something; couldn't you also validate your model on the test set? That will only allow you to validate the counts, not the predicted coordinates of the animals, but wouldn't that be enough?",
      "votes": null
    },
    {
      "id": "182914",
      "postDate": "05/16/2017 09:47:48",
      "content": "<p>@erde not sure I understand what you mean - I can't use the test set for validation because we don't have the counts for the test set - we need to predict them.</p>",
      "rawMarkdown": "erde not sure I understand what you mean - I can't use the test set for validation because we don't have the counts for the test set - we need to predict them.",
      "votes": null
    },
    {
      "id": "182988",
      "postDate": "05/16/2017 16:19:24",
      "content": "<p>I think that the discrepancy between local validation score and public leaderboard may arise from difference in calculation of RMSE. The metric for the competition is average columnwise RMSE, the metric used for example by Keras is rowwise RMSE. </p>\n\n<p>For a 3x5 table \n[1  2  3  4  5]\n[6  7  8  9  10]\n[11 12 13 14 15]\nwith ground truth of 0</p>\n\n<p>RMSE columnwise = 9.1\nRMS rowwise = 15.8</p>\n\n<p>On the bright side they move in the same direction, so if your rowwise RMSE is going down, that means your columnwise RMSE will go down too.</p>",
      "rawMarkdown": "I think that the discrepancy between local validation score and public leaderboard may arise from difference in calculation of RMSE. The metric for the competition is average columnwise RMSE, the metric used for example by Keras is rowwise RMSE. \n\nFor a 3x5 table \n[1  2  3  4  5]\n[6  7  8  9  10]\n[11 12 13 14 15]\nwith ground truth of 0\n\nRMSE columnwise = 9.1\nRMS rowwise = 15.8\n\nOn the bright side they move in the same direction, so if your rowwise RMSE is going down, that means your columnwise RMSE will go down too.",
      "votes": null
    },
    {
      "id": "182990",
      "postDate": "05/16/2017 16:29:58",
      "content": "<p>This gap could be explained by the metric. RMSE between different datasets can only be compared if the average counts are similar. In case you did the train/validation split blindly you could end up with a validation set with lower counts (when counts are lower the same model scores better).</p>",
      "rawMarkdown": "This gap could be explained by the metric. RMSE between different datasets can only be compared if the average counts are similar. In case you did the train/validation split blindly you could end up with a validation set with lower counts (when counts are lower the same model scores better).",
      "votes": null
    },
    {
      "id": "182994",
      "postDate": "05/16/2017 16:46:42",
      "content": "<p>Yes, maybe this is the reason, because I'm getting completely different values.\nFor your example (assuming you meant columns of length 3 and rows of length 5), my mean column-wise RMSE is 9.0052, and my mean row-wise RMSE is 8.1724.</p>\n\n<p>Can someone else try calculating it? :)</p>",
      "rawMarkdown": "Yes, maybe this is the reason, because I'm getting completely different values.\nFor your example (assuming you meant columns of length 3 and rows of length 5), my mean column-wise RMSE is 9.0052, and my mean row-wise RMSE is 8.1724.\n\nCan someone else try calculating it? :)",
      "votes": null
    },
    {
      "id": "182997",
      "postDate": "05/16/2017 16:57:45",
      "content": "<p>Yes, some difference (e.g. scale) between train and test set is my leading hypothesis at the moment, assuming I implemented RMSE calculation correctly. I don't think it's explained just by a bad split though.</p>",
      "rawMarkdown": "Yes, some difference (e.g. scale) between train and test set is my leading hypothesis at the moment, assuming I implemented RMSE calculation correctly. I don't think it's explained just by a bad split though.",
      "votes": null
    },
    {
      "id": "183136",
      "postDate": "05/17/2017 02:58:46",
      "content": "<p>I got a huge gap, too. I think row-wise RMSE is lower than column-wise in this competition.\nYou can check your submission's column-wise RMS (RMSE to all zero) vs. 29 (public LB RMS).\nThe public LB score always worse than abs(submission's RMS - 29).\nIf model isn't good enough, the public LB score will worse than abs(submission's RMS - 29) +  local validation score.</p>",
      "rawMarkdown": "I got a huge gap, too. I think row-wise RMSE is lower than column-wise in this competition.\nYou can check your submission's column-wise RMS (RMSE to all zero) vs. 29 (public LB RMS).\nThe public LB score always worse than abs(submission's RMS - 29).\nIf model isn't good enough, the public LB score will worse than abs(submission's RMS - 29) +  local validation score.",
      "votes": null
    },
    {
      "id": "183150",
      "postDate": "05/17/2017 05:16:35",
      "content": "<p>Thanks for the clarification. I misunderstood the structure of the dataset. But you could automate uploading the test set predictions to Kaggle to get a score, or would that evaluation be too noisy/slow?</p>",
      "rawMarkdown": "Thanks for the clarification. I misunderstood the structure of the dataset. But you could automate uploading the test set predictions to Kaggle to get a score, or would that evaluation be too noisy/slow?",
      "votes": null
    },
    {
      "id": "183278",
      "postDate": "05/17/2017 16:20:58",
      "content": "<p>@erde that's very dangerous, as you can overfit to the public LB and then fail on the private LB. So it's better to have a working local CV. Also, you can have only 5 submissions per day, which is not a lot.</p>",
      "rawMarkdown": "erde that's very dangerous, as you can overfit to the public LB and then fail on the private LB. So it's better to have a working local CV. Also, you can have only 5 submissions per day, which is not a lot.",
      "votes": null
    },
    {
      "id": "183442",
      "postDate": "05/18/2017 08:54:47",
      "content": "<p>Public and private LB are different subsets of the files in <code>input/Test</code>?</p>",
      "rawMarkdown": "Public and private LB are different subsets of the files in `input/Test`?",
      "votes": null
    },
    {
      "id": "183446",
      "postDate": "05/18/2017 09:14:09",
      "content": "<p>Same values for me: 9.00520 column-wise, 8.17245 row-wise.</p>",
      "rawMarkdown": "Same values for me: 9.00520 column-wise, 8.17245 row-wise.",
      "votes": null
    },
    {
      "id": "183578",
      "postDate": "05/18/2017 16:41:22",
      "content": "<p>whew, thanks Lowik!</p>",
      "rawMarkdown": "whew, thanks Lowik!",
      "votes": null
    },
    {
      "id": "183580",
      "postDate": "05/18/2017 16:50:14",
      "content": "<p>&gt; Public and private LB are different subsets of the files in input/Test?</p>\n\n<p>Right, if you go to the Leaderboard page, it says: \"This leaderboard is calculated with approximately 50% of the test data. The final results will be based on the other 50%, so the final standings may be different.\"</p>\n\n<p>Besides that, only a small portion of the \"input/Test\" is labeled (no one knows, but probably on the order of 1000 images), so probably the public LB is about 500-1000 images, and the private LB is  the same.</p>",
      "rawMarkdown": "&gt; Public and private LB are different subsets of the files in input/Test?\n\nRight, if you go to the Leaderboard page, it says: \"This leaderboard is calculated with approximately 50% of the test data. The final results will be based on the other 50%, so the final standings may be different.\"\n\nBesides that, only a small portion of the \"input/Test\" is labeled (no one knows, but probably on the order of 1000 images), so probably the public LB is about 500-1000 images, and the private LB is  the same.",
      "votes": null
    },
    {
      "id": "189006",
      "postDate": "06/04/2017 20:02:00",
      "content": "<p>I have some problems calculating RMSE:</p>\n\n<p>First of all seems new version of Keras doesn't have RMSE (only MSE):</p>\n\n<p>But it's not hard to implement:</p>\n\n<pre><code>def mean_squared_error(y_true, y_pred):\n    \"\"\"\n    MSE loss function\n    \"\"\"\n    return K.mean(K.square(y_pred - y_true), axis=-1)\n\ndef root_mean_squared_error(y_true, y_pred):\n    \"\"\"\n    RMSE loss function\n    \"\"\"\n    return K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n</code></pre>\n\n<p>Than I have tried to calculate it by hand and it doesn't the same as keras output:</p>\n\n<pre><code>def  compute_RMSE(y_true, y_pred):\n    RMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n\n    return RMSE\n\ndef evalute_on_val(model,X_val,y_val):\n    y_pred= model.predict(X_val)\n    RMSE= compute_RMSE(y_val, y_pred)\n\n    print 'RMSE:', RMSE\n\n\nmodel.fit(X_train, Y_train, validation_data=(X_val, Y_val), batch_size=batch_size, epochs=1)\nprint model.evaluate(X_val, Y_val)\nevalute_on_val(model, X_val, Y_val)\n</code></pre>\n\n<blockquote>\n  <p>Epoch 1/1 32/32 [==============================] - 4s - loss: 4.0545 -\n  val_loss: 1.9832 224/228 [============================&gt;.] - ETA:\n  0s1.98315653048 RMSE: 2.14734170033 Train on 32 samples, validate on\n  228 samples Epoch 1/1 32/32 [==============================] - 2s -\n  loss: 3.0376 - val_loss: 1.9279 224/228\n  [============================&gt;.] - ETA: 0s1.92787208892 RMSE:\n  2.08691443128</p>\n</blockquote>\n\n<p>So it's 1.98 vs 2.08</p>\n\n<p>Also here is seems related issue: <a href=\"https://github.com/fchollet/keras/issues/1170\">https://github.com/fchollet/keras/issues/1170</a></p>",
      "rawMarkdown": "I have some problems calculating RMSE:\n\nFirst of all seems new version of Keras doesn't have RMSE (only MSE):\n\nBut it's not hard to implement:\n\n    def mean_squared_error(y_true, y_pred):\n    \t\"\"\"\n    \tMSE loss function\n    \t\"\"\"\n    \treturn K.mean(K.square(y_pred - y_true), axis=-1)\n    \n    def root_mean_squared_error(y_true, y_pred):\n    \t\"\"\"\n    \tRMSE loss function\n    \t\"\"\"\n    \treturn K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n\nThan I have tried to calculate it by hand and it doesn't the same as keras output:\n\n    def  compute_RMSE(y_true, y_pred):\n    \tRMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n    \t\n    \treturn RMSE\n    \n    def evalute_on_val(model,X_val,y_val):\n    \ty_pred= model.predict(X_val)\n    \tRMSE= compute_RMSE(y_val, y_pred)\n    \n    \tprint 'RMSE:', RMSE\n\n\n    model.fit(X_train, Y_train, validation_data=(X_val, Y_val), batch_size=batch_size, epochs=1)\n    print model.evaluate(X_val, Y_val)\n    evalute_on_val(model, X_val, Y_val)\n\n&gt; Epoch 1/1 32/32 [==============================] - 4s - loss: 4.0545 -\n&gt; val_loss: 1.9832 224/228 [============================&gt;.] - ETA:\n&gt; 0s1.98315653048 RMSE: 2.14734170033 Train on 32 samples, validate on\n&gt; 228 samples Epoch 1/1 32/32 [==============================] - 2s -\n&gt; loss: 3.0376 - val_loss: 1.9279 224/228\n&gt; [============================&gt;.] - ETA: 0s1.92787208892 RMSE:\n&gt; 2.08691443128\n\nSo it's 1.98 vs 2.08\n\nAlso here is seems related issue: https://github.com/fchollet/keras/issues/1170",
      "votes": null
    },
    {
      "id": "192393",
      "postDate": "06/13/2017 12:08:36",
      "content": "<p>BTW Leaderboard is 'Mean Columnwise Root Mean Squared Error'</p>",
      "rawMarkdown": "BTW Leaderboard is 'Mean Columnwise Root Mean Squared Error'",
      "votes": null
    },
    {
      "id": "192513",
      "postDate": "06/13/2017 20:53:29",
      "content": "<p>Hmm, but I still wonder what is row-wise rmse?</p>\n\n<pre><code>#This is like in sklearn\ndef compute_rmse(y_true, y_pred):\n    RMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n\n    return RMSE\n\ndef test_rmse():\n    y_arr= np.array([[1,2,3,4,5],[6,7,8,9,10],[11,12,13,14,15]])\n\n    print 'RMSE:', compute_rmse(y_arr, np.zeros(y_arr.shape))\n\n    from sklearn.metrics import mean_squared_error\n    print 'RMSE sklearn:', mean_squared_error(y_arr, np.zeros(y_arr.shape))**0.5\n</code></pre>\n\n<blockquote>\n  <p>RMSE: 9.09212113132 RMSE sklearn: 9.09212113132</p>\n</blockquote>",
      "rawMarkdown": "Hmm, but I still wonder what is row-wise rmse?\n\n    #This is like in sklearn\n    def compute_rmse(y_true, y_pred):\n    \tRMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n    \t\n    \treturn RMSE\n    \n    def test_rmse():\n    \ty_arr= np.array([[1,2,3,4,5],[6,7,8,9,10],[11,12,13,14,15]])\n    \n    \tprint 'RMSE:', compute_rmse(y_arr, np.zeros(y_arr.shape))\n    \t\n    \tfrom sklearn.metrics import mean_squared_error\n    \tprint 'RMSE sklearn:', mean_squared_error(y_arr, np.zeros(y_arr.shape))**0.5\n\n&gt; RMSE: 9.09212113132 RMSE sklearn: 9.09212113132",
      "votes": null
    },
    {
      "id": "195734",
      "postDate": "06/24/2017 19:43:10",
      "content": "<p>By \"local\" you refer what exactly? The corrected SeaLionData tid's with the masks from TrainDotted applied to the images from Train?</p>",
      "rawMarkdown": "By \"local\" you refer what exactly? The corrected SeaLionData tid's with the masks from TrainDotted applied to the images from Train?",
      "votes": null
    },
    {
      "id": "195744",
      "postDate": "06/24/2017 21:17:04",
      "content": "<blockquote>\n  <p>By \"local\" you refer what exactly?</p>\n</blockquote>\n\n<p>Run against validation set that I constructed from a part of Train. Here \"local\" means local as opposed to the public leaderboard which is run on the Kaggle servers.</p>",
      "rawMarkdown": "&gt; By \"local\" you refer what exactly?\n\nRun against validation set that I constructed from a part of Train. Here \"local\" means local as opposed to the public leaderboard which is run on the Kaggle servers.",
      "votes": null
    },
    {
      "id": "195745",
      "postDate": "06/24/2017 21:20:05",
      "content": "<p>I see, Thanks! But I guess you did add the masks to them from TrainDotted? By the way, what RMSE do you now get on the local validation set?</p>",
      "rawMarkdown": "I see, Thanks! But I guess you did add the masks to them from TrainDotted? By the way, what RMSE do you now get on the local validation set?",
      "votes": null
    },
    {
      "id": "195838",
      "postDate": "06/25/2017 08:21:32",
      "content": "<blockquote>\n  <p>But I guess you did add the masks to them from TrainDotted?</p>\n</blockquote>\n\n<p>Yes, for training</p>\n\n<blockquote>\n  <p>By the way, what RMSE do you now get on the local validation set?</p>\n</blockquote>\n\n<p>10 -- 16</p>",
      "rawMarkdown": "&gt; But I guess you did add the masks to them from TrainDotted?\n\nYes, for training\n\n&gt; By the way, what RMSE do you now get on the local validation set?\n\n10 -- 16",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 179651,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "05/02/2017 12:15:08",
      "content": "<p>Konstantin, do you train a CNN on first ~750 images and than tune classifiers on hold-out 250 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 179657,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/02/2017 12:27:35",
          "content": "<p>Yes - actually, I'm training a CNN on about 650 images, and use about 100 for validation of the CNN, and then tune regression models that take CNN predictions as input on the hold-out 250.</p>\n\n<p>Edit: and I'm also excluding images from MismatchedTrainImages.txt + 941,  200, 491, 912</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 182912,
          "author_name": "traces",
          "author_url": "",
          "post_date": "05/16/2017 09:44:20",
          "content": "<p>Just to understand your approach, and I apologize in advance if I misunderstood something; couldn't you also validate your model on the test set? That will only allow you to validate the counts, not the predicted coordinates of the animals, but wouldn't that be enough?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 182914,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/16/2017 09:47:48",
          "content": "<p>@erde not sure I understand what you mean - I can't use the test set for validation because we don't have the counts for the test set - we need to predict them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183150,
          "author_name": "traces",
          "author_url": "",
          "post_date": "05/17/2017 05:16:35",
          "content": "<p>Thanks for the clarification. I misunderstood the structure of the dataset. But you could automate uploading the test set predictions to Kaggle to get a score, or would that evaluation be too noisy/slow?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183278,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/17/2017 16:20:58",
          "content": "<p>@erde that's very dangerous, as you can overfit to the public LB and then fail on the private LB. So it's better to have a working local CV. Also, you can have only 5 submissions per day, which is not a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183442,
          "author_name": "traces",
          "author_url": "",
          "post_date": "05/18/2017 08:54:47",
          "content": "<p>Public and private LB are different subsets of the files in <code>input/Test</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183580,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/18/2017 16:50:14",
          "content": "<p>&gt; Public and private LB are different subsets of the files in input/Test?</p>\n\n<p>Right, if you go to the Leaderboard page, it says: \"This leaderboard is calculated with approximately 50% of the test data. The final results will be based on the other 50%, so the final standings may be different.\"</p>\n\n<p>Besides that, only a small portion of the \"input/Test\" is labeled (no one knows, but probably on the order of 1000 images), so probably the public LB is about 500-1000 images, and the private LB is  the same.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 179688,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "05/02/2017 14:11:23",
      "content": "<p>My current approach make massive gap too.<br>\nIn my opinion, we should not believe public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 179767,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/02/2017 19:49:49",
          "content": "<blockquote>\n  <p>In my opinion, we should not believe public LB.</p>\n</blockquote>\n\n<p>Do you think that the private LB will be better? Do you have any hypothesis about that gap between local CV and public LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 179906,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "05/03/2017 08:14:47",
          "content": "<p>Some competitors point out that number of sea-lions in <code>train.csv</code> don't consistent real number. This implies <code>train.csv</code>　includes much noise. I don't know that private LB will be better, but I hope our final submission evaluated with real number.<br></p>\n\n<p>By the way, admin shared <code>MismatchedTrainImages.txt</code> after starting competition. I wonder if there is <code>MismatchedTestImages.txt</code> potentially or not.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 182904,
      "author_name": "hagorms",
      "author_url": "",
      "post_date": "05/16/2017 09:00:58",
      "content": "<p>Hello.\nHave the same problem =(\nI tried to submit all zeros, and it gave me the score 29.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 182988,
      "author_name": "mitiau",
      "author_url": "",
      "post_date": "05/16/2017 16:19:24",
      "content": "<p>I think that the discrepancy between local validation score and public leaderboard may arise from difference in calculation of RMSE. The metric for the competition is average columnwise RMSE, the metric used for example by Keras is rowwise RMSE. </p>\n\n<p>For a 3x5 table \n[1  2  3  4  5]\n[6  7  8  9  10]\n[11 12 13 14 15]\nwith ground truth of 0</p>\n\n<p>RMSE columnwise = 9.1\nRMS rowwise = 15.8</p>\n\n<p>On the bright side they move in the same direction, so if your rowwise RMSE is going down, that means your columnwise RMSE will go down too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 182994,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/16/2017 16:46:42",
          "content": "<p>Yes, maybe this is the reason, because I'm getting completely different values.\nFor your example (assuming you meant columns of length 3 and rows of length 5), my mean column-wise RMSE is 9.0052, and my mean row-wise RMSE is 8.1724.</p>\n\n<p>Can someone else try calculating it? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183446,
          "author_name": "nzeuwik",
          "author_url": "",
          "post_date": "05/18/2017 09:14:09",
          "content": "<p>Same values for me: 9.00520 column-wise, 8.17245 row-wise.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 183578,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/18/2017 16:41:22",
          "content": "<p>whew, thanks Lowik!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 192393,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/13/2017 12:08:36",
          "content": "<p>BTW Leaderboard is 'Mean Columnwise Root Mean Squared Error'</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 192513,
          "author_name": "mrgloom",
          "author_url": "",
          "post_date": "06/13/2017 20:53:29",
          "content": "<p>Hmm, but I still wonder what is row-wise rmse?</p>\n\n<pre><code>#This is like in sklearn\ndef compute_rmse(y_true, y_pred):\n    RMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n\n    return RMSE\n\ndef test_rmse():\n    y_arr= np.array([[1,2,3,4,5],[6,7,8,9,10],[11,12,13,14,15]])\n\n    print 'RMSE:', compute_rmse(y_arr, np.zeros(y_arr.shape))\n\n    from sklearn.metrics import mean_squared_error\n    print 'RMSE sklearn:', mean_squared_error(y_arr, np.zeros(y_arr.shape))**0.5\n</code></pre>\n\n<blockquote>\n  <p>RMSE: 9.09212113132 RMSE sklearn: 9.09212113132</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 182990,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "05/16/2017 16:29:58",
      "content": "<p>This gap could be explained by the metric. RMSE between different datasets can only be compared if the average counts are similar. In case you did the train/validation split blindly you could end up with a validation set with lower counts (when counts are lower the same model scores better).</p>",
      "votes": null,
      "replies": [
        {
          "id": 182997,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "05/16/2017 16:57:45",
          "content": "<p>Yes, some difference (e.g. scale) between train and test set is my leading hypothesis at the moment, assuming I implemented RMSE calculation correctly. I don't think it's explained just by a bad split though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 183136,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "05/17/2017 02:58:46",
      "content": "<p>I got a huge gap, too. I think row-wise RMSE is lower than column-wise in this competition.\nYou can check your submission's column-wise RMS (RMSE to all zero) vs. 29 (public LB RMS).\nThe public LB score always worse than abs(submission's RMS - 29).\nIf model isn't good enough, the public LB score will worse than abs(submission's RMS - 29) +  local validation score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 189006,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/04/2017 20:02:00",
      "content": "<p>I have some problems calculating RMSE:</p>\n\n<p>First of all seems new version of Keras doesn't have RMSE (only MSE):</p>\n\n<p>But it's not hard to implement:</p>\n\n<pre><code>def mean_squared_error(y_true, y_pred):\n    \"\"\"\n    MSE loss function\n    \"\"\"\n    return K.mean(K.square(y_pred - y_true), axis=-1)\n\ndef root_mean_squared_error(y_true, y_pred):\n    \"\"\"\n    RMSE loss function\n    \"\"\"\n    return K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n</code></pre>\n\n<p>Than I have tried to calculate it by hand and it doesn't the same as keras output:</p>\n\n<pre><code>def  compute_RMSE(y_true, y_pred):\n    RMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n\n    return RMSE\n\ndef evalute_on_val(model,X_val,y_val):\n    y_pred= model.predict(X_val)\n    RMSE= compute_RMSE(y_val, y_pred)\n\n    print 'RMSE:', RMSE\n\n\nmodel.fit(X_train, Y_train, validation_data=(X_val, Y_val), batch_size=batch_size, epochs=1)\nprint model.evaluate(X_val, Y_val)\nevalute_on_val(model, X_val, Y_val)\n</code></pre>\n\n<blockquote>\n  <p>Epoch 1/1 32/32 [==============================] - 4s - loss: 4.0545 -\n  val_loss: 1.9832 224/228 [============================&gt;.] - ETA:\n  0s1.98315653048 RMSE: 2.14734170033 Train on 32 samples, validate on\n  228 samples Epoch 1/1 32/32 [==============================] - 2s -\n  loss: 3.0376 - val_loss: 1.9279 224/228\n  [============================&gt;.] - ETA: 0s1.92787208892 RMSE:\n  2.08691443128</p>\n</blockquote>\n\n<p>So it's 1.98 vs 2.08</p>\n\n<p>Also here is seems related issue: <a href=\"https://github.com/fchollet/keras/issues/1170\">https://github.com/fchollet/keras/issues/1170</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195734,
      "author_name": "traces",
      "author_url": "",
      "post_date": "06/24/2017 19:43:10",
      "content": "<p>By \"local\" you refer what exactly? The corrected SeaLionData tid's with the masks from TrainDotted applied to the images from Train?</p>",
      "votes": null,
      "replies": [
        {
          "id": 195744,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/24/2017 21:17:04",
          "content": "<blockquote>\n  <p>By \"local\" you refer what exactly?</p>\n</blockquote>\n\n<p>Run against validation set that I constructed from a part of Train. Here \"local\" means local as opposed to the public leaderboard which is run on the Kaggle servers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195745,
          "author_name": "traces",
          "author_url": "",
          "post_date": "06/24/2017 21:20:05",
          "content": "<p>I see, Thanks! But I guess you did add the masks to them from TrainDotted? By the way, what RMSE do you now get on the local validation set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195838,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "06/25/2017 08:21:32",
          "content": "<blockquote>\n  <p>But I guess you did add the masks to them from TrainDotted?</p>\n</blockquote>\n\n<p>Yes, for training</p>\n\n<blockquote>\n  <p>By the way, what RMSE do you now get on the local validation set?</p>\n</blockquote>\n\n<p>10 -- 16</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "179595": "Has anyone built a reliable local cross-validation strategy? For me the gap is massive, up to 12 on local vs 24 on LB.\n\nI'm leaving about 250 random images for the model that predicts counts (so the CNN does not see them at all). Then I'm doing a 5-fold cross-validation on this 250 images using a couple of models such as XGBoost and Lasso.\n\nWhat could explain such a huge gap? It could obviously be a bug in my code, or very different images in test vs. train (I didn't check predictions on test but the images look similar). Apart from that, I don't see any obvious reasons, or ways I'm leaking train into my validation.",
    "179651": "Konstantin, do you train a CNN on first ~750 images and than tune classifiers on hold-out 250 ?",
    "179657": "Yes - actually, I'm training a CNN on about 650 images, and use about 100 for validation of the CNN, and then tune regression models that take CNN predictions as input on the hold-out 250.\n\nEdit: and I'm also excluding images from MismatchedTrainImages.txt + 941,  200, 491, 912",
    "179688": "My current approach make massive gap too.<br>\nIn my opinion, we should not believe public LB.",
    "179767": "> In my opinion, we should not believe public LB.\n\nDo you think that the private LB will be better? Do you have any hypothesis about that gap between local CV and public LB?",
    "179906": "Some competitors point out that number of sea-lions in `train.csv` don't consistent real number. This implies `train.csv`　includes much noise. I don't know that private LB will be better, but I hope our final submission evaluated with real number.<br>\n\nBy the way, admin shared `MismatchedTrainImages.txt` after starting competition. I wonder if there is `MismatchedTestImages.txt` potentially or not.",
    "182904": "Hello.\nHave the same problem =(\nI tried to submit all zeros, and it gave me the score 29.",
    "182912": "Just to understand your approach, and I apologize in advance if I misunderstood something; couldn't you also validate your model on the test set? That will only allow you to validate the counts, not the predicted coordinates of the animals, but wouldn't that be enough?",
    "182914": "erde not sure I understand what you mean - I can't use the test set for validation because we don't have the counts for the test set - we need to predict them.",
    "182988": "I think that the discrepancy between local validation score and public leaderboard may arise from difference in calculation of RMSE. The metric for the competition is average columnwise RMSE, the metric used for example by Keras is rowwise RMSE. \n\nFor a 3x5 table \n[1  2  3  4  5]\n[6  7  8  9  10]\n[11 12 13 14 15]\nwith ground truth of 0\n\nRMSE columnwise = 9.1\nRMS rowwise = 15.8\n\nOn the bright side they move in the same direction, so if your rowwise RMSE is going down, that means your columnwise RMSE will go down too.",
    "182990": "This gap could be explained by the metric. RMSE between different datasets can only be compared if the average counts are similar. In case you did the train/validation split blindly you could end up with a validation set with lower counts (when counts are lower the same model scores better).",
    "182994": "Yes, maybe this is the reason, because I'm getting completely different values.\nFor your example (assuming you meant columns of length 3 and rows of length 5), my mean column-wise RMSE is 9.0052, and my mean row-wise RMSE is 8.1724.\n\nCan someone else try calculating it? :)",
    "182997": "Yes, some difference (e.g. scale) between train and test set is my leading hypothesis at the moment, assuming I implemented RMSE calculation correctly. I don't think it's explained just by a bad split though.",
    "183136": "I got a huge gap, too. I think row-wise RMSE is lower than column-wise in this competition.\nYou can check your submission's column-wise RMS (RMSE to all zero) vs. 29 (public LB RMS).\nThe public LB score always worse than abs(submission's RMS - 29).\nIf model isn't good enough, the public LB score will worse than abs(submission's RMS - 29) +  local validation score.",
    "183150": "Thanks for the clarification. I misunderstood the structure of the dataset. But you could automate uploading the test set predictions to Kaggle to get a score, or would that evaluation be too noisy/slow?",
    "183278": "erde that's very dangerous, as you can overfit to the public LB and then fail on the private LB. So it's better to have a working local CV. Also, you can have only 5 submissions per day, which is not a lot.",
    "183442": "Public and private LB are different subsets of the files in `input/Test`?",
    "183446": "Same values for me: 9.00520 column-wise, 8.17245 row-wise.",
    "183578": "whew, thanks Lowik!",
    "183580": "&gt; Public and private LB are different subsets of the files in input/Test?\n\nRight, if you go to the Leaderboard page, it says: \"This leaderboard is calculated with approximately 50% of the test data. The final results will be based on the other 50%, so the final standings may be different.\"\n\nBesides that, only a small portion of the \"input/Test\" is labeled (no one knows, but probably on the order of 1000 images), so probably the public LB is about 500-1000 images, and the private LB is  the same.",
    "189006": "I have some problems calculating RMSE:\n\nFirst of all seems new version of Keras doesn't have RMSE (only MSE):\n\nBut it's not hard to implement:\n\n    def mean_squared_error(y_true, y_pred):\n    \t\"\"\"\n    \tMSE loss function\n    \t\"\"\"\n    \treturn K.mean(K.square(y_pred - y_true), axis=-1)\n    \n    def root_mean_squared_error(y_true, y_pred):\n    \t\"\"\"\n    \tRMSE loss function\n    \t\"\"\"\n    \treturn K.sqrt(K.mean(K.square(y_pred - y_true), axis=-1))\n\nThan I have tried to calculate it by hand and it doesn't the same as keras output:\n\n    def  compute_RMSE(y_true, y_pred):\n    \tRMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n    \t\n    \treturn RMSE\n    \n    def evalute_on_val(model,X_val,y_val):\n    \ty_pred= model.predict(X_val)\n    \tRMSE= compute_RMSE(y_val, y_pred)\n    \n    \tprint 'RMSE:', RMSE\n\n\n    model.fit(X_train, Y_train, validation_data=(X_val, Y_val), batch_size=batch_size, epochs=1)\n    print model.evaluate(X_val, Y_val)\n    evalute_on_val(model, X_val, Y_val)\n\n&gt; Epoch 1/1 32/32 [==============================] - 4s - loss: 4.0545 -\n&gt; val_loss: 1.9832 224/228 [============================&gt;.] - ETA:\n&gt; 0s1.98315653048 RMSE: 2.14734170033 Train on 32 samples, validate on\n&gt; 228 samples Epoch 1/1 32/32 [==============================] - 2s -\n&gt; loss: 3.0376 - val_loss: 1.9279 224/228\n&gt; [============================&gt;.] - ETA: 0s1.92787208892 RMSE:\n&gt; 2.08691443128\n\nSo it's 1.98 vs 2.08\n\nAlso here is seems related issue: https://github.com/fchollet/keras/issues/1170",
    "192393": "BTW Leaderboard is 'Mean Columnwise Root Mean Squared Error'",
    "192513": "Hmm, but I still wonder what is row-wise rmse?\n\n    #This is like in sklearn\n    def compute_rmse(y_true, y_pred):\n    \tRMSE= np.sqrt(np.mean((y_true-y_pred)**2))\n    \t\n    \treturn RMSE\n    \n    def test_rmse():\n    \ty_arr= np.array([[1,2,3,4,5],[6,7,8,9,10],[11,12,13,14,15]])\n    \n    \tprint 'RMSE:', compute_rmse(y_arr, np.zeros(y_arr.shape))\n    \t\n    \tfrom sklearn.metrics import mean_squared_error\n    \tprint 'RMSE sklearn:', mean_squared_error(y_arr, np.zeros(y_arr.shape))**0.5\n\n&gt; RMSE: 9.09212113132 RMSE sklearn: 9.09212113132",
    "195734": "By \"local\" you refer what exactly? The corrected SeaLionData tid's with the masks from TrainDotted applied to the images from Train?",
    "195744": "&gt; By \"local\" you refer what exactly?\n\nRun against validation set that I constructed from a part of Train. Here \"local\" means local as opposed to the public leaderboard which is run on the Kaggle servers.",
    "195745": "I see, Thanks! But I guess you did add the masks to them from TrainDotted? By the way, what RMSE do you now get on the local validation set?",
    "195838": "&gt; But I guess you did add the masks to them from TrainDotted?\n\nYes, for training\n\n&gt; By the way, what RMSE do you now get on the local validation set?\n\n10 -- 16"
  },
  "source": "meta"
}