{
  "id": 89722,
  "title": "How do you determine your CV score?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89722",
  "author_name": "",
  "post_date": "2019-04-17T03:11:21.659344900Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Let's say you run a 5-fold CV. So in each of the 5 folds, there will be a CV (MAE) score. So my question is - how will you calculate your CV score?</p>\n\n<p>Option 1: you take the average (and std) of the 5 folds of the cv score. \nOption 2: you save the prediction of testing data from each fold, aggregate them as a single OOF vector, and calculate the overall MAE score against the training target after the cross-validation such as\n<code>\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n</code></p>",
  "messages": [
    {
      "id": "518271",
      "postDate": "04/17/2019 03:11:21",
      "content": "<p>Let's say you run a 5-fold CV. So in each of the 5 folds, there will be a CV (MAE) score. So my question is - how will you calculate your CV score?</p>\n\n<p>Option 1: you take the average (and std) of the 5 folds of the cv score. \nOption 2: you save the prediction of testing data from each fold, aggregate them as a single OOF vector, and calculate the overall MAE score against the training target after the cross-validation such as\n<code>\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n</code></p>",
      "rawMarkdown": "Let's say you run a 5-fold CV. So in each of the 5 folds, there will be a CV (MAE) score. So my question is - how will you calculate your CV score?\n\nOption 1: you take the average (and std) of the 5 folds of the cv score. \nOption 2: you save the prediction of testing data from each fold, aggregate them as a single OOF vector, and calculate the overall MAE score against the training target after the cross-validation such as\n```\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n```",
      "votes": null
    },
    {
      "id": "518286",
      "postDate": "04/17/2019 03:41:27",
      "content": "<p>Usually people follow the first option.</p>",
      "rawMarkdown": "Usually people follow the first option.",
      "votes": null
    },
    {
      "id": "518431",
      "postDate": "04/17/2019 07:44:26",
      "content": "<p>I think option2 is better,option1 is a little overfitting.</p>",
      "rawMarkdown": "I think option2 is better,option1 is a little overfitting.",
      "votes": null
    },
    {
      "id": "518454",
      "postDate": "04/17/2019 08:09:52",
      "content": "<p>Option 3: stratified CV&lt;--&gt;taking 3 (approx) earthquakes in a validation sample. With suffling = less deviation but usually lower MAE. Speaking locall CV ofcourse....</p>",
      "rawMarkdown": "Option 3: stratified CV&lt;--&gt;taking 3 (approx) earthquakes in a validation sample. With suffling = less deviation but usually lower MAE. Speaking locall CV ofcourse....",
      "votes": null
    },
    {
      "id": "518497",
      "postDate": "04/17/2019 11:34:51",
      "content": "<p>Currently, I am using option 2 because if you split the CV by earthquake groups, the MAE for the first and last earthquake will be much better than the rest and hence unreasonably lower the MAE when compared with taking the average (option 1). I am not sure which one is more trustworthy though.</p>",
      "rawMarkdown": "Currently, I am using option 2 because if you split the CV by earthquake groups, the MAE for the first and last earthquake will be much better than the rest and hence unreasonably lower the MAE when compared with taking the average (option 1). I am not sure which one is more trustworthy though.",
      "votes": null
    },
    {
      "id": "518657",
      "postDate": "04/17/2019 15:19:14",
      "content": "<p><code>\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n</code></p>\n\n<p>That is how I do it</p>",
      "rawMarkdown": "```\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n```\n\nThat is how I do it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 518286,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "04/17/2019 03:41:27",
      "content": "<p>Usually people follow the first option.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518431,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "04/17/2019 07:44:26",
      "content": "<p>I think option2 is better,option1 is a little overfitting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518454,
      "author_name": "zikazika",
      "author_url": "",
      "post_date": "04/17/2019 08:09:52",
      "content": "<p>Option 3: stratified CV&lt;--&gt;taking 3 (approx) earthquakes in a validation sample. With suffling = less deviation but usually lower MAE. Speaking locall CV ofcourse....</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518497,
      "author_name": "pukkinming",
      "author_url": "",
      "post_date": "04/17/2019 11:34:51",
      "content": "<p>Currently, I am using option 2 because if you split the CV by earthquake groups, the MAE for the first and last earthquake will be much better than the rest and hence unreasonably lower the MAE when compared with taking the average (option 1). I am not sure which one is more trustworthy though.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 518657,
      "author_name": "halldalton94",
      "author_url": "",
      "post_date": "04/17/2019 15:19:14",
      "content": "<p><code>\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n</code></p>\n\n<p>That is how I do it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "518271": "Let's say you run a 5-fold CV. So in each of the 5 folds, there will be a CV (MAE) score. So my question is - how will you calculate your CV score?\n\nOption 1: you take the average (and std) of the 5 folds of the cv score. \nOption 2: you save the prediction of testing data from each fold, aggregate them as a single OOF vector, and calculate the overall MAE score against the training target after the cross-validation such as\n```\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n```",
    "518286": "Usually people follow the first option.",
    "518431": "I think option2 is better,option1 is a little overfitting.",
    "518454": "Option 3: stratified CV&lt;--&gt;taking 3 (approx) earthquakes in a validation sample. With suffling = less deviation but usually lower MAE. Speaking locall CV ofcourse....",
    "518497": "Currently, I am using option 2 because if you split the CV by earthquake groups, the MAE for the first and last earthquake will be much better than the rest and hence unreasonably lower the MAE when compared with taking the average (option 1). I am not sure which one is more trustworthy though.",
    "518657": "```\nfor fold_n, (train_index, valid_index) in enumerate(folds):\n   .....\n   oof[valid_index] = y_pred_valid.reshape(-1,)\ncv_score = mean_absolute_error(oof, y)\n```\n\nThat is how I do it"
  },
  "source": "meta"
}