{
  "id": 83328,
  "title": "train CV scores are worse?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/83328",
  "author_name": "",
  "post_date": "2019-03-08T16:58:47.802422300Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm wondering if I calculate the scoring correctly - since all my train scores (with single sample set, or with n-fold CV) are always worse than my kaggle scores.\nTrain is always around 2.00 to 2.10 while the same model scores 1.50 / 1.60 when uploaded to kaggle.</p>\n\n<p>I use:\nmean(abs(predictionTTF-testTTF))</p>\n\n<p>Am I wrong on the scoring?</p>",
  "messages": [
    {
      "id": "486358",
      "postDate": "03/08/2019 16:58:47",
      "content": "<p>I'm wondering if I calculate the scoring correctly - since all my train scores (with single sample set, or with n-fold CV) are always worse than my kaggle scores.\nTrain is always around 2.00 to 2.10 while the same model scores 1.50 / 1.60 when uploaded to kaggle.</p>\n\n<p>I use:\nmean(abs(predictionTTF-testTTF))</p>\n\n<p>Am I wrong on the scoring?</p>",
      "rawMarkdown": "I'm wondering if I calculate the scoring correctly - since all my train scores (with single sample set, or with n-fold CV) are always worse than my kaggle scores.\nTrain is always around 2.00 to 2.10 while the same model scores 1.50 / 1.60 when uploaded to kaggle.\n\nI use:\nmean(abs(predictionTTF-testTTF))\n\nAm I wrong on the scoring?",
      "votes": null
    },
    {
      "id": "486547",
      "postDate": "03/08/2019 23:57:11",
      "content": "<p>me too.. puzzling </p>",
      "rawMarkdown": "me too.. puzzling",
      "votes": null
    },
    {
      "id": "487203",
      "postDate": "03/10/2019 10:52:13",
      "content": "<p>This can be explained by the small quantity of samples used in the set used to calculate the public score (only 13% of the whole test set). </p>\n\n<p>The public set is probably not statisticaly significant, and it is probably better to try to optimize your cross-validation score than the public score.</p>",
      "rawMarkdown": "This can be explained by the small quantity of samples used in the set used to calculate the public score (only 13% of the whole test set). \n\nThe public set is probably not statisticaly significant, and it is probably better to try to optimize your cross-validation score than the public score.",
      "votes": null
    },
    {
      "id": "487684",
      "postDate": "03/11/2019 10:27:03",
      "content": "<p>Hmmm, it still boggles my mind a bit.\nSurely this must mean that our training sets contains more extremes,\nand the LB test-set is really average, so training data scores really well on it?\nThis might also mean that training (and CV'ing) the set without the extremes might give a better result? </p>\n\n<p>For example leaving out TTF&lt;0.5 sec and TTF&gt;9 sec. \nI'm grasping to understand :) - thanks for the reply anyways</p>",
      "rawMarkdown": "Hmmm, it still boggles my mind a bit.\nSurely this must mean that our training sets contains more extremes,\nand the LB test-set is really average, so training data scores really well on it?\nThis might also mean that training (and CV'ing) the set without the extremes might give a better result? \n\nFor example leaving out TTF&lt;0.5 sec and TTF&gt;9 sec. \nI'm grasping to understand :) - thanks for the reply anyways",
      "votes": null
    },
    {
      "id": "488035",
      "postDate": "03/11/2019 20:28:48",
      "content": "<p>Indeed, I'm with @Jacky here. The testing dataset is small and probably not statistically significant. My guess is that in the testing dataset there are not large values of <code>time_to_failure</code>.</p>",
      "rawMarkdown": "Indeed, I'm with @Jacky here. The testing dataset is small and probably not statistically significant. My guess is that in the testing dataset there are not large values of `time_to_failure`.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 486547,
      "author_name": "asterisk",
      "author_url": "",
      "post_date": "03/08/2019 23:57:11",
      "content": "<p>me too.. puzzling </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 487203,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "03/10/2019 10:52:13",
      "content": "<p>This can be explained by the small quantity of samples used in the set used to calculate the public score (only 13% of the whole test set). </p>\n\n<p>The public set is probably not statisticaly significant, and it is probably better to try to optimize your cross-validation score than the public score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 487684,
          "author_name": "theupgrade",
          "author_url": "",
          "post_date": "03/11/2019 10:27:03",
          "content": "<p>Hmmm, it still boggles my mind a bit.\nSurely this must mean that our training sets contains more extremes,\nand the LB test-set is really average, so training data scores really well on it?\nThis might also mean that training (and CV'ing) the set without the extremes might give a better result? </p>\n\n<p>For example leaving out TTF&lt;0.5 sec and TTF&gt;9 sec. \nI'm grasping to understand :) - thanks for the reply anyways</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 488035,
      "author_name": "ricarddelgado",
      "author_url": "",
      "post_date": "03/11/2019 20:28:48",
      "content": "<p>Indeed, I'm with @Jacky here. The testing dataset is small and probably not statistically significant. My guess is that in the testing dataset there are not large values of <code>time_to_failure</code>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "486358": "I'm wondering if I calculate the scoring correctly - since all my train scores (with single sample set, or with n-fold CV) are always worse than my kaggle scores.\nTrain is always around 2.00 to 2.10 while the same model scores 1.50 / 1.60 when uploaded to kaggle.\n\nI use:\nmean(abs(predictionTTF-testTTF))\n\nAm I wrong on the scoring?",
    "486547": "me too.. puzzling",
    "487203": "This can be explained by the small quantity of samples used in the set used to calculate the public score (only 13% of the whole test set). \n\nThe public set is probably not statisticaly significant, and it is probably better to try to optimize your cross-validation score than the public score.",
    "487684": "Hmmm, it still boggles my mind a bit.\nSurely this must mean that our training sets contains more extremes,\nand the LB test-set is really average, so training data scores really well on it?\nThis might also mean that training (and CV'ing) the set without the extremes might give a better result? \n\nFor example leaving out TTF&lt;0.5 sec and TTF&gt;9 sec. \nI'm grasping to understand :) - thanks for the reply anyways",
    "488035": "Indeed, I'm with @Jacky here. The testing dataset is small and probably not statistically significant. My guess is that in the testing dataset there are not large values of `time_to_failure`."
  },
  "source": "meta"
}