{
  "id": 10405,
  "title": "Python code for CV splitting",
  "url": "/competitions/seizure-prediction/discussion/10405",
  "author_name": "",
  "post_date": "2014-09-21T22:13:54.077Z",
  "votes": 6,
  "comment_count": 18,
  "views": 5697,
  "content": "<p><a href=\"https://github.com/ebenolson/seizure-prediction-public\">https://github.com/ebenolson/seizure-prediction-public</a></p>\n<p>Thought I'd share some code I wrote to split&nbsp;the data&nbsp;into test/train sets, while keeping sequences intact - this&nbsp;should help&nbsp;avoid&nbsp;overly-optimistic CV scores.</p>\n<p>Hope it's useful, and please let me know if you notice&nbsp;any mistakes.</p>",
  "messages": [
    {
      "id": "54390",
      "postDate": "09/21/2014 22:13:54",
      "content": "<p><a href=\"https://github.com/ebenolson/seizure-prediction-public\">https://github.com/ebenolson/seizure-prediction-public</a></p>\n<p>Thought I'd share some code I wrote to split&nbsp;the data&nbsp;into test/train sets, while keeping sequences intact - this&nbsp;should help&nbsp;avoid&nbsp;overly-optimistic CV scores.</p>\n<p>Hope it's useful, and please let me know if you notice&nbsp;any mistakes.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54393",
      "postDate": "09/22/2014 01:32:25",
      "content": "<p>I'm sorry , could you explain the benefit of keeping sequences intact ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54395",
      "postDate": "09/22/2014 01:57:48",
      "content": "<p>Data segments within the same sequence will be more similar to each other than to other sequences, so if you just choose a random split your classifier will have an easier job&nbsp;and you may overestimate its performance.</p>\n<p>Splitting by sequence should give more accurate estimates, although for some subjects there are so few sequences it will be quite&nbsp;noisy.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "54396",
      "postDate": "09/22/2014 02:02:20",
      "content": "<p>Oh ok, i think i misinterpreted what you were doing. Thanks for the code !.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55553",
      "postDate": "10/03/2014 15:02:19",
      "content": "<p>If you use Python, I would suggest you using scikit-learn for cross-validation. If you have 4-core CPU, it automatically parallelize the 4-fold CV and you run 4 times faster, for free.</p>\n<p>Hand crafting code is respectful, but not very productive.</p>\n<p>http://scikit-learn.org/stable/modules/cross_validation.html</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55555",
      "postDate": "10/03/2014 15:07:47",
      "content": "<p>Is there a method in sklearn that let's you preserve a ratio between positive and negative samples in the train split ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55557",
      "postDate": "10/03/2014 15:24:19",
      "content": "<p>Fortunately&nbsp;there is.</p>\n<p>http://scikit-learn.org/dev/modules/generated/sklearn.cross_validation.StratifiedShuffleSplit.html#sklearn.cross_validation.StratifiedShuffleSplit</p>\n<p>http://scikit-learn.org/dev/modules/generated/sklearn.cross_validation.StratifiedKFold.html#sklearn.cross_validation.StratifiedKFold</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "55569",
      "postDate": "10/03/2014 19:32:23",
      "content": "<p>thanks !</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56171",
      "postDate": "10/17/2014 09:22:14",
      "content": "<p>Hi, I'm looking at the provided code and it seems to me that it's the equivalent of using two calls to train_test_split in scikit (one for preictal and interictal). Can someone confirm?</p>\n<p>I'm also curious what sort of differences you guys are seeing between the leaderboard scores and your CV scores using this method.</p>\n\n<p>http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html#sklearn.cross_validation.train_test_split</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56173",
      "postDate": "10/17/2014 10:08:42",
      "content": "<p>[quote=Nissan Pow;56171]</p>\n<p>Hi, I'm looking at the provided code and it seems to me that it's the equivalent of using two calls to train_test_split in scikit (one for preictal and interictal). Can someone confirm?</p>\n<p>I'm also curious what sort of differences you guys are seeing between the leaderboard scores and your CV scores using this method.</p>\n<p>http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html#sklearn.cross_validation.train_test_split</p>\n<p>[/quote]</p>\n<p>No. Using a simple random split you can easily get a CV score around 0.96 (which is totally inaccurate).&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56174",
      "postDate": "10/17/2014 11:36:17",
      "content": "<p>[quote=emolson;56173]</p>\n<p>[quote=Nissan Pow;56171]</p>\n<p>Hi, I'm looking at the provided code and it seems to me that it's the equivalent of using two calls to train_test_split in scikit (one for preictal and interictal). Can someone confirm?</p>\n<p>I'm also curious what sort of differences you guys are seeing between the leaderboard scores and your CV scores using this method.</p>\n<p>http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html#sklearn.cross_validation.train_test_split</p>\n<p>[/quote]</p>\n<p>No. Using a simple random split you can easily get a CV score around 0.96 (which is totally inaccurate).&nbsp;</p>\n<p>[/quote]</p>\n<p>Sorry, I meant splitting the preictals and interictals separately within each subject (not just a random split&nbsp;over all the training data). I see that for each subject you're splitting the preictal and interictal separately&nbsp;with shuffling.</p>\n<p>The reason I'm asking is that I was getting ridiculous CV scores using shuffling. However they're more realistic without, which is contrary to your findings.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56176",
      "postDate": "10/17/2014 12:11:02",
      "content": "<p>Oh, I see what you mean.</p>\n<p>As you probably realize, the issue is that clips within a sequence are extremely similar and it is easy to classify clips if you have a training example from the same sequence.</p>\n<p>Splitting without shuffling is preferable to shuffling and splitting, but it will have the same problem to a lesser extent - since not all the sequences&nbsp;have the same number of clips, your split will likely divide one of them.</p>\n<p>The point of this code is to make sure that the test and train sets are made up of independent complete&nbsp;sequences.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "56177",
      "postDate": "10/17/2014 13:04:10",
      "content": "<p>[quote=emolson;56176]</p>\n<p>Oh, I see what you mean.</p>\n<p>As you probably realize, the issue is that clips within a sequence are extremely similar and it is easy to classify clips if you have a training example from the same sequence.</p>\n<p>Splitting without shuffling is preferable to shuffling and splitting, but it will have the same problem to a lesser extent - since not all the sequences&nbsp;have the same number of clips, your split will likely divide one of them.</p>\n<p>The point of this code is to make sure that the test and train sets are made up of independent complete&nbsp;sequences.</p>\n<p>[/quote]</p>\n\n<p>Thanks for the clarification! I also just realized that&nbsp;your code has all the sequences grouped in the pickle file, which is how it works (it's really&nbsp;subject name -&gt; type -&gt; list of <strong>list of filenames</strong>).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57017",
      "postDate": "10/29/2014 21:23:03",
      "content": "<p>This is really great. Just to be clear here, though, from what I can tell, the sequences of items come in groups of 6, correct?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57021",
      "postDate": "10/29/2014 21:42:17",
      "content": "<p>Unfortunately it's not quite that simple, as a number of the sequences don't have all&nbsp;6 clips - I don't remember but I think some only have 1 or 2.&nbsp;</p>\n<p>The grouping data is in 'filenames.pickle'</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57035",
      "postDate": "10/30/2014 08:09:55",
      "content": "<p>In case anyone is interested (or for anyone not using python), here is a csv file with segment names and sequence that were extracted from filenames.pickle.</p>\n<p>(Ignore the file with '+' in the filename. Apparently the upload doesn't like files named that way...)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57240",
      "postDate": "11/03/2014 13:44:24",
      "content": "<p>[quote=Maineiac;57035]</p>\n<p>In case anyone is interested (or for anyone not using python), here is a csv file with segment names and sequence that were extracted from filenames.pickle.</p>\n<p>(Ignore the file with '+' in the filename. Apparently the upload doesn't like files named that way...)</p>\n<p>[/quote]</p>\n<p>Hi, thanks for providing the list, why you have 480 Dog_5 and I have 671?</p>\n\n\n<p>Oh, I see, that's for training.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57562",
      "postDate": "11/09/2014 18:27:00",
      "content": "<p>Dumb question but can someone clarify for me in plain English what exactly a sequence and a segment are?</p>\n\n<p>Here's my current understanding:&nbsp;</p>\n<p>Segments from the same sequence occurred close together in time. Different sequences represent different time periods. True?</p>\n<p>Is there a seizure in each sequence?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "57567",
      "postDate": "11/09/2014 18:53:20",
      "content": "<p>[quote=Thomas O'Malley;57562]</p>\n<p>Dumb question but can someone clarify for me in plain English what exactly a sequence and a segment are?</p>\n<p>Here's my current understanding:&nbsp;</p>\n<p>Segments from the same sequence occurred close together in time. Different sequences represent different time periods. True?</p>\n<p>Is there a seizure in each sequence?</p>\n<p>[/quote]</p>\n\n<p>No,there is a seizure 10mins after each sequence.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 54393,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/22/2014 01:32:25",
      "content": "<p>I'm sorry , could you explain the benefit of keeping sequences intact ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54395,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "09/22/2014 01:57:48",
      "content": "<p>Data segments within the same sequence will be more similar to each other than to other sequences, so if you just choose a random split your classifier will have an easier job&nbsp;and you may overestimate its performance.</p>\n<p>Splitting by sequence should give more accurate estimates, although for some subjects there are so few sequences it will be quite&nbsp;noisy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 54396,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "09/22/2014 02:02:20",
      "content": "<p>Oh ok, i think i misinterpreted what you were doing. Thanks for the code !.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55553,
      "author_name": "",
      "author_url": "",
      "post_date": "10/03/2014 15:02:19",
      "content": "<p>If you use Python, I would suggest you using scikit-learn for cross-validation. If you have 4-core CPU, it automatically parallelize the 4-fold CV and you run 4 times faster, for free.</p>\n<p>Hand crafting code is respectful, but not very productive.</p>\n<p>http://scikit-learn.org/stable/modules/cross_validation.html</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55555,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "10/03/2014 15:07:47",
      "content": "<p>Is there a method in sklearn that let's you preserve a ratio between positive and negative samples in the train split ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55557,
      "author_name": "",
      "author_url": "",
      "post_date": "10/03/2014 15:24:19",
      "content": "<p>Fortunately&nbsp;there is.</p>\n<p>http://scikit-learn.org/dev/modules/generated/sklearn.cross_validation.StratifiedShuffleSplit.html#sklearn.cross_validation.StratifiedShuffleSplit</p>\n<p>http://scikit-learn.org/dev/modules/generated/sklearn.cross_validation.StratifiedKFold.html#sklearn.cross_validation.StratifiedKFold</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 55569,
      "author_name": "franklyn",
      "author_url": "",
      "post_date": "10/03/2014 19:32:23",
      "content": "<p>thanks !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56171,
      "author_name": "nissanpow",
      "author_url": "",
      "post_date": "10/17/2014 09:22:14",
      "content": "<p>Hi, I'm looking at the provided code and it seems to me that it's the equivalent of using two calls to train_test_split in scikit (one for preictal and interictal). Can someone confirm?</p>\n<p>I'm also curious what sort of differences you guys are seeing between the leaderboard scores and your CV scores using this method.</p>\n\n<p>http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html#sklearn.cross_validation.train_test_split</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56173,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "10/17/2014 10:08:42",
      "content": "<p>[quote=Nissan Pow;56171]</p>\n<p>Hi, I'm looking at the provided code and it seems to me that it's the equivalent of using two calls to train_test_split in scikit (one for preictal and interictal). Can someone confirm?</p>\n<p>I'm also curious what sort of differences you guys are seeing between the leaderboard scores and your CV scores using this method.</p>\n<p>http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html#sklearn.cross_validation.train_test_split</p>\n<p>[/quote]</p>\n<p>No. Using a simple random split you can easily get a CV score around 0.96 (which is totally inaccurate).&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56174,
      "author_name": "nissanpow",
      "author_url": "",
      "post_date": "10/17/2014 11:36:17",
      "content": "<p>[quote=emolson;56173]</p>\n<p>[quote=Nissan Pow;56171]</p>\n<p>Hi, I'm looking at the provided code and it seems to me that it's the equivalent of using two calls to train_test_split in scikit (one for preictal and interictal). Can someone confirm?</p>\n<p>I'm also curious what sort of differences you guys are seeing between the leaderboard scores and your CV scores using this method.</p>\n<p>http://scikit-learn.org/stable/modules/generated/sklearn.cross_validation.train_test_split.html#sklearn.cross_validation.train_test_split</p>\n<p>[/quote]</p>\n<p>No. Using a simple random split you can easily get a CV score around 0.96 (which is totally inaccurate).&nbsp;</p>\n<p>[/quote]</p>\n<p>Sorry, I meant splitting the preictals and interictals separately within each subject (not just a random split&nbsp;over all the training data). I see that for each subject you're splitting the preictal and interictal separately&nbsp;with shuffling.</p>\n<p>The reason I'm asking is that I was getting ridiculous CV scores using shuffling. However they're more realistic without, which is contrary to your findings.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56176,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "10/17/2014 12:11:02",
      "content": "<p>Oh, I see what you mean.</p>\n<p>As you probably realize, the issue is that clips within a sequence are extremely similar and it is easy to classify clips if you have a training example from the same sequence.</p>\n<p>Splitting without shuffling is preferable to shuffling and splitting, but it will have the same problem to a lesser extent - since not all the sequences&nbsp;have the same number of clips, your split will likely divide one of them.</p>\n<p>The point of this code is to make sure that the test and train sets are made up of independent complete&nbsp;sequences.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 56177,
      "author_name": "nissanpow",
      "author_url": "",
      "post_date": "10/17/2014 13:04:10",
      "content": "<p>[quote=emolson;56176]</p>\n<p>Oh, I see what you mean.</p>\n<p>As you probably realize, the issue is that clips within a sequence are extremely similar and it is easy to classify clips if you have a training example from the same sequence.</p>\n<p>Splitting without shuffling is preferable to shuffling and splitting, but it will have the same problem to a lesser extent - since not all the sequences&nbsp;have the same number of clips, your split will likely divide one of them.</p>\n<p>The point of this code is to make sure that the test and train sets are made up of independent complete&nbsp;sequences.</p>\n<p>[/quote]</p>\n\n<p>Thanks for the clarification! I also just realized that&nbsp;your code has all the sequences grouped in the pickle file, which is how it works (it's really&nbsp;subject name -&gt; type -&gt; list of <strong>list of filenames</strong>).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57017,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "10/29/2014 21:23:03",
      "content": "<p>This is really great. Just to be clear here, though, from what I can tell, the sequences of items come in groups of 6, correct?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57021,
      "author_name": "emolson",
      "author_url": "",
      "post_date": "10/29/2014 21:42:17",
      "content": "<p>Unfortunately it's not quite that simple, as a number of the sequences don't have all&nbsp;6 clips - I don't remember but I think some only have 1 or 2.&nbsp;</p>\n<p>The grouping data is in 'filenames.pickle'</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57035,
      "author_name": "maineiac",
      "author_url": "",
      "post_date": "10/30/2014 08:09:55",
      "content": "<p>In case anyone is interested (or for anyone not using python), here is a csv file with segment names and sequence that were extracted from filenames.pickle.</p>\n<p>(Ignore the file with '+' in the filename. Apparently the upload doesn't like files named that way...)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57240,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/03/2014 13:44:24",
      "content": "<p>[quote=Maineiac;57035]</p>\n<p>In case anyone is interested (or for anyone not using python), here is a csv file with segment names and sequence that were extracted from filenames.pickle.</p>\n<p>(Ignore the file with '+' in the filename. Apparently the upload doesn't like files named that way...)</p>\n<p>[/quote]</p>\n<p>Hi, thanks for providing the list, why you have 480 Dog_5 and I have 671?</p>\n\n\n<p>Oh, I see, that's for training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57562,
      "author_name": "omalleyt",
      "author_url": "",
      "post_date": "11/09/2014 18:27:00",
      "content": "<p>Dumb question but can someone clarify for me in plain English what exactly a sequence and a segment are?</p>\n\n<p>Here's my current understanding:&nbsp;</p>\n<p>Segments from the same sequence occurred close together in time. Different sequences represent different time periods. True?</p>\n<p>Is there a seizure in each sequence?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 57567,
      "author_name": "stevendu",
      "author_url": "",
      "post_date": "11/09/2014 18:53:20",
      "content": "<p>[quote=Thomas O'Malley;57562]</p>\n<p>Dumb question but can someone clarify for me in plain English what exactly a sequence and a segment are?</p>\n<p>Here's my current understanding:&nbsp;</p>\n<p>Segments from the same sequence occurred close together in time. Different sequences represent different time periods. True?</p>\n<p>Is there a seizure in each sequence?</p>\n<p>[/quote]</p>\n\n<p>No,there is a seizure 10mins after each sequence.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "54390": "",
    "54393": "",
    "54395": "",
    "54396": "",
    "55553": "",
    "55555": "",
    "55557": "",
    "55569": "",
    "56171": "",
    "56173": "",
    "56174": "",
    "56176": "",
    "56177": "",
    "57017": "",
    "57021": "",
    "57035": "",
    "57240": "",
    "57562": "",
    "57567": ""
  },
  "source": "meta"
}