{
  "id": 89579,
  "title": "Best Practices for CV",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89579",
  "author_name": "",
  "post_date": "2019-04-16T01:28:44.867150500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Is there a consensus for the best method to cross validate? Does simple Kfolds do the trick, or should individual earthquakes not be separated?</p>",
  "messages": [
    {
      "id": "517388",
      "postDate": "04/16/2019 01:28:44",
      "content": "<p>Is there a consensus for the best method to cross validate? Does simple Kfolds do the trick, or should individual earthquakes not be separated?</p>",
      "rawMarkdown": "Is there a consensus for the best method to cross validate? Does simple Kfolds do the trick, or should individual earthquakes not be separated?",
      "votes": null
    },
    {
      "id": "517969",
      "postDate": "04/16/2019 17:56:44",
      "content": "<p>Not sure that there's a consensus, but separating by earthquake and doing leave-one-earthquake-out CV might not be the best given that this <a href=\"https://www.kaggle.com/miklgr500/fast-failure-detector\">kernel</a> shows that failure may occur in the middle of the test segments. There are only 7 of these in the test set though, so it probably won't impact the score too much.</p>",
      "rawMarkdown": "Not sure that there's a consensus, but separating by earthquake and doing leave-one-earthquake-out CV might not be the best given that this [kernel](https://www.kaggle.com/miklgr500/fast-failure-detector) shows that failure may occur in the middle of the test segments. There are only 7 of these in the test set though, so it probably won't impact the score too much.",
      "votes": null
    },
    {
      "id": "518441",
      "postDate": "04/17/2019 07:52:24",
      "content": "<p>I think Machine Learning as an Experimental Science,you should try more things.</p>",
      "rawMarkdown": "I think Machine Learning as an Experimental Science,you should try more things.",
      "votes": null
    },
    {
      "id": "518670",
      "postDate": "04/17/2019 15:33:14",
      "content": "<p>I agree, at this point it is up to me to try different validation strategies.</p>",
      "rawMarkdown": "I agree, at this point it is up to me to try different validation strategies.",
      "votes": null
    },
    {
      "id": "518770",
      "postDate": "04/17/2019 19:05:06",
      "content": "<p>The failures shown in this kernel are <em>within</em> the same quake cycle. You can check it on the train set: you will see big spikes at low ttf, but the spike is not the boundary between one cycle and the other. In this respect if you separate the quake cycles in the CV I don't see any problem </p>",
      "rawMarkdown": "The failures shown in this kernel are *within* the same quake cycle. You can check it on the train set: you will see big spikes at low ttf, but the spike is not the boundary between one cycle and the other. In this respect if you separate the quake cycles in the CV I don't see any problem",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 517969,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "04/16/2019 17:56:44",
      "content": "<p>Not sure that there's a consensus, but separating by earthquake and doing leave-one-earthquake-out CV might not be the best given that this <a href=\"https://www.kaggle.com/miklgr500/fast-failure-detector\">kernel</a> shows that failure may occur in the middle of the test segments. There are only 7 of these in the test set though, so it probably won't impact the score too much.</p>",
      "votes": null,
      "replies": [
        {
          "id": 518770,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "04/17/2019 19:05:06",
          "content": "<p>The failures shown in this kernel are <em>within</em> the same quake cycle. You can check it on the train set: you will see big spikes at low ttf, but the spike is not the boundary between one cycle and the other. In this respect if you separate the quake cycles in the CV I don't see any problem </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 518441,
      "author_name": "senkin13",
      "author_url": "",
      "post_date": "04/17/2019 07:52:24",
      "content": "<p>I think Machine Learning as an Experimental Science,you should try more things.</p>",
      "votes": null,
      "replies": [
        {
          "id": 518670,
          "author_name": "halldalton94",
          "author_url": "",
          "post_date": "04/17/2019 15:33:14",
          "content": "<p>I agree, at this point it is up to me to try different validation strategies.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "517388": "Is there a consensus for the best method to cross validate? Does simple Kfolds do the trick, or should individual earthquakes not be separated?",
    "517969": "Not sure that there's a consensus, but separating by earthquake and doing leave-one-earthquake-out CV might not be the best given that this [kernel](https://www.kaggle.com/miklgr500/fast-failure-detector) shows that failure may occur in the middle of the test segments. There are only 7 of these in the test set though, so it probably won't impact the score too much.",
    "518441": "I think Machine Learning as an Experimental Science,you should try more things.",
    "518670": "I agree, at this point it is up to me to try different validation strategies.",
    "518770": "The failures shown in this kernel are *within* the same quake cycle. You can check it on the train set: you will see big spikes at low ttf, but the spike is not the boundary between one cycle and the other. In this respect if you separate the quake cycles in the CV I don't see any problem"
  },
  "source": "meta"
}