{
  "id": 9544,
  "title": "Need some help understanding stacked generalisation.",
  "url": "/competitions/decoding-the-human-brain/discussion/9544",
  "author_name": "",
  "post_date": "2014-06-24T20:50:15.823Z",
  "votes": null,
  "comment_count": 10,
  "views": 6034,
  "content": "<p>Hi all,</p>\n<p>I will appreciate some help on understanding stacked generalization. I understand SG applied to a statistically coherent data set, with candidate classifiers being a variety of algorithms and models. However, I am a bit confused on its application to this data set. I am seeing two alternatives:</p>\n<p>Alternative 1: There are 16 classifiers, each say a logistic regression.<br>Step 1: Partition the training data into per subject data set, TD1..TD16.<br>Step2: Fit one classifier exclusively on one subject. Thus clf1 is trained on TD1 ..and so on.<br>Step3: Generate probability estimates on entire data set, for each classifier. <br>Step 4: Train a single level 1 classifier on these probability estimates using same labels as train data.</p>\n<p>When a test vector is exposed to the 16 level 0 classifiers, they generate a probability vector that resembles a level of class membership for each subject. The level 1 classifier further classifies this to a face/scramble. This approach makes sense. However, I see a performance identical to pooling , no improvement.</p>\n<p>Alternative 2: This is motivated by the statement in this paper and elsewhere, promoting 'cross-training'&nbsp;across training data.e.g. the statement &#8220;predicted values of a given train comes from classifier which were not trained on that trail&#8221;</p>\n<p><br>There are 16 classifiers, each say a logistic regression.<br>Step 1: Partition the training data into per subject data set, TD1..TD16.<br>Step2: Fit one classifier all train data EXCLUDING one subject. Thus clf1 is trained on TD2 thru TD16, excluding TD1 and so on.<br>Step3: Generate probability estimates on entire data set, for each classifier. <br>Step 4: Train a single level 1 classifier on these probability estimates using same labels as train data.</p>\n<p>When a test vector is exposed to the 16 level 0 classifiers, they generate a probability vector that resembles a biased level of class membership for each subject. In that, if a test vector were to be drawn from the same statistics as subject used for training, the class membership for that subject will not stand out. This does not seem correct and my LB score validates this assumption.</p>\n<p>If approach 1 is correct, I will go ahead and debug my algorithm, particularly the preprocessing. But at this time there does not seem to be a coding error. Any pointers please?</p>",
  "messages": [
    {
      "id": "49504",
      "postDate": "06/24/2014 20:50:15",
      "content": "<p>Hi all,</p>\n<p>I will appreciate some help on understanding stacked generalization. I understand SG applied to a statistically coherent data set, with candidate classifiers being a variety of algorithms and models. However, I am a bit confused on its application to this data set. I am seeing two alternatives:</p>\n<p>Alternative 1: There are 16 classifiers, each say a logistic regression.<br>Step 1: Partition the training data into per subject data set, TD1..TD16.<br>Step2: Fit one classifier exclusively on one subject. Thus clf1 is trained on TD1 ..and so on.<br>Step3: Generate probability estimates on entire data set, for each classifier. <br>Step 4: Train a single level 1 classifier on these probability estimates using same labels as train data.</p>\n<p>When a test vector is exposed to the 16 level 0 classifiers, they generate a probability vector that resembles a level of class membership for each subject. The level 1 classifier further classifies this to a face/scramble. This approach makes sense. However, I see a performance identical to pooling , no improvement.</p>\n<p>Alternative 2: This is motivated by the statement in this paper and elsewhere, promoting 'cross-training'&nbsp;across training data.e.g. the statement &#8220;predicted values of a given train comes from classifier which were not trained on that trail&#8221;</p>\n<p><br>There are 16 classifiers, each say a logistic regression.<br>Step 1: Partition the training data into per subject data set, TD1..TD16.<br>Step2: Fit one classifier all train data EXCLUDING one subject. Thus clf1 is trained on TD2 thru TD16, excluding TD1 and so on.<br>Step3: Generate probability estimates on entire data set, for each classifier. <br>Step 4: Train a single level 1 classifier on these probability estimates using same labels as train data.</p>\n<p>When a test vector is exposed to the 16 level 0 classifiers, they generate a probability vector that resembles a biased level of class membership for each subject. In that, if a test vector were to be drawn from the same statistics as subject used for training, the class membership for that subject will not stand out. This does not seem correct and my LB score validates this assumption.</p>\n<p>If approach 1 is correct, I will go ahead and debug my algorithm, particularly the preprocessing. But at this time there does not seem to be a coding error. Any pointers please?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49520",
      "postDate": "06/25/2014 02:20:41",
      "content": "<p>My results got worse when I tried to used stacked generalization as described in Alternative 1. I have not tried Alternative 2.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49534",
      "postDate": "06/25/2014 05:58:04",
      "content": "<p>Trent,</p>\n<p>When I do a cross validation over say 30% of training data, with alternative 1, with the 30% test data drawn uniformly from all subjects, i do see an improvement over pooling and&nbsp;is explicable. When I do a leave-one out cross validation, i do not see any improvement over pooling, which while aligning with LB scores, seems to suggest a missing ingredient somewhere, given that multiple people have reported seeing improvements.</p>\n\n<p>Given your significantly higher score, and the fact that you are not using SG, I am curious how you see such an improvement. Is it due to preprocessing, and/or covariate shift only? Are you doing some other kind of ensemble method?</p>\n<p>Thanks</p>\n<p>Kalpendu</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49595",
      "postDate": "06/26/2014 00:36:52",
      "content": "<p>I spent most of my time on preprocessing to get features that generalized well before trying stacked generalization or a method to handle covariate shift.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "49742",
      "postDate": "06/27/2014 23:09:49",
      "content": "<p>Trent,</p>\n<p>Thanks for the confirmation.</p>\n<p>I think I understand what is going on-my intuition is that SG will yield maximum benefit in case of minimally&nbsp;processed signal. I am using the fourier space as basis for pre processing, which I believe is destroying ability to discriminate the subjects. Fourier is extremely&nbsp;sub-optimal for random signal. The feature it generates yields SG useless. My feeling is it will impact covariate shift as well. No data to substantiate, will update if i confirm&nbsp;anything.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "50059",
      "postDate": "07/05/2014 06:25:48",
      "content": "<p>I'm having trouble visualizing the flow of data from one level to the next in ensemble learning/stack generalization. Admittedly, this is the first time I'm working with this concept, so please bear with me as I haven't fully wrapped my head around it.&nbsp;</p>\n\n<p>I spent some time today going through DH Wolpert's (1992) paper on stacked generalization, and reading through Olivetti et. al. (2014) which is posted on the main page of this competition. I start to get tripped up when I read through the following portion on page 6 of Olivetti's paper:</p>\n\n<ol>\n<li>Train a set of classi&#64257;ers on (portions of) the train data. These classi&#64257;ers&nbsp;are called &#64257;rst-level classi&#64257;ers.</li>\n<li>Collect the output of each classi&#64257;er on each trial of the train and of the&nbsp;test data. These outputs are called &#64257;rst level predictions.</li>\n<li><strong>Create a new second-level dataset with the vector of &#64257;rst-level predictions&nbsp;</strong><strong>for each trial. Care has to be taken so that the predicted value of a&nbsp;</strong><strong>given trial comes from classi&#64257;ers which were not trained on that trial, e.g.&nbsp;</strong><strong>through cross-validation.</strong></li>\n<li>The class-labels of the second-level dataset are the same as the initial&nbsp;dataset.</li>\n<li>A second-level classi&#64257;er is trained on the portion of the second-level dataset&nbsp;related to the train subjects in order to learn how to combine the &#64257;rst-level&nbsp;predictions.</li>\n<li>The second level classi&#64257;er is used to predict the class-labels of the test&nbsp;data as represented in the second-level dataset.</li>\n</ol>\n<p>Based on my interpretation, I would&nbsp;expect the second-level dataset to contain the same number of examples as the original train data (we'll call it m), and one feature for each classifier trained on each one&nbsp;of the 16 subjects. Then again, that would seem to betray the requirement that &quot;[C]are has to be taken so that the predicted value of a&nbsp;given trial comes from classi&#64257;ers which were not trained on that trial, e.g.&nbsp;through cross-validation.&quot; For instance, classifier #1 would have been built using all of the trials associated with subject #1. Not sure how to get around this issue.</p>\n\n<p>I've searched around for every paper I could find today on ensemble learning and stacked generalization and have not been able to figure this out (looked at the two listed above plus Sigletos et. al. (2005) and Ting et. al (1999)). If anyone has any advice it would be very helpful.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "50060",
      "postDate": "07/05/2014 07:23:39",
      "content": "<p>I had the same confusion. My interpretation&nbsp;is:</p>\n<p>1. For each first level, subject specific classifier k, k=1..16 do the following:</p>\n<p>a .&nbsp;Divide &nbsp;data for each subject, k, into J data sets</p>\n<p>b. create J CV classifier.</p>\n<p>c. Train jth classifier using the rest j-1 data sets.</p>\n<p>d.Create L1 meta data, i.e. 'probabilities' for the data set 'j'&nbsp;&nbsp;using the jth classifier. This resolves the issue you mention</p>\n<p>e. Create a 'master' classifier for the kth subject. This classifier is trained using all the subject specific data.</p>\n<p>f. Use the L1 meta-data in step d to&nbsp;as training set for a L1 classifier</p>\n\n<p>2. For a new test data, run it through the master classifiers for each subject.</p>\n<p>3. Use the probalilities as meta-data to classify using the L1 classifier.</p>\n\n<p>I see a marginal improvement with this.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "50064",
      "postDate": "07/05/2014 15:06:32",
      "content": "<p>So let me make sure that I understand. Let's say that we have a data set of 3 subjects each with 3 trials for a total of 9 trials.&nbsp;</p>\n\n<p>First, we would focus specifically on subject #1. Within subject 1, we would first focus on trial #1. &nbsp;We would use the other two trials for subject #1 (ie trials 2 and 3) &nbsp;to create a classifier, and then we would use this classifier to predict trial #1.&nbsp;</p>\n\n<p>Then, we would move on to trial #2, where we would use trials 1 and 3 within subject #1 to make a classifier, and then use that classifier to predict #2.&nbsp;</p>\n\n<p>We we would repeat this process one more time for trial #3., using 1 and 2 to build the classifier and making a prediction for trial #3.</p>\n\n<p>At this point, we are done with subject #1 and should have 3 predictions, one for each of subject 1's trials.</p>\n\n<p>We would then follow this process for subjects 2 and 3. &nbsp;Ultimately, we end up with a vector of m &nbsp;= 9 predictions, each of which used the other trials within the same subject to build the classifier.</p>\n\n<p>Then we could use this vector to train a new classifier, using the predictions from the preceding step as inputs, and using the original classes as outputs. &nbsp;</p>\n\n<p>Is this correct, or did I misunderstand?</p>\n\n<p>Thanks very much for your help!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "50065",
      "postDate": "07/05/2014 15:14:41",
      "content": "<p>Although I'm not sure I understand the piece on building a 'master classifier' for each subject..</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "50071",
      "postDate": "07/06/2014 05:48:22",
      "content": "<p>Rob,</p>\n<p>Correct understanding. What classifier model would you use to generate L1 meta data on a new test data? You need one single model. This is the master classifier.&nbsp;</p>\n<p>Disclaimer-this is my understanding and i do not have any concrete data to prove that this is correct.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "51270",
      "postDate": "07/28/2014 17:34:00",
      "content": "<p>I implemented stack generalization with some mixed results. Although our final result was better with a pooled model. Here is what I did,</p>\n\n<p>1) Each subject had about 600 trials, I split that into sets of 500 and 100. I trained a classifier on each of the subjects using just the 500 trials.&nbsp;So now I had 16 classifiers. </p>\n<p>2) I used the remaining 100 trials from each to make predictions on each of the 16 classifiers. So this gives me 100*16 trials for my second level classifier. </p>\n<p>3) I ran the whole test set through each of the first level classifiers, used the predictions on the second level classifier.</p>\n<p>This gave me mixed results. Some subjects bumped in accuracy by as much as 0.63-0.74</p>\n<p>Certain subjects like subject 3,16 were pretty hard to predict..they went worse. One the leaderboard this scheme was giving me just about 0.67...so we abandoned it but this was definitely worth trying more.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 49520,
      "author_name": "trentb",
      "author_url": "",
      "post_date": "06/25/2014 02:20:41",
      "content": "<p>My results got worse when I tried to used stacked generalization as described in Alternative 1. I have not tried Alternative 2.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49534,
      "author_name": "wabbit",
      "author_url": "",
      "post_date": "06/25/2014 05:58:04",
      "content": "<p>Trent,</p>\n<p>When I do a cross validation over say 30% of training data, with alternative 1, with the 30% test data drawn uniformly from all subjects, i do see an improvement over pooling and&nbsp;is explicable. When I do a leave-one out cross validation, i do not see any improvement over pooling, which while aligning with LB scores, seems to suggest a missing ingredient somewhere, given that multiple people have reported seeing improvements.</p>\n\n<p>Given your significantly higher score, and the fact that you are not using SG, I am curious how you see such an improvement. Is it due to preprocessing, and/or covariate shift only? Are you doing some other kind of ensemble method?</p>\n<p>Thanks</p>\n<p>Kalpendu</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49595,
      "author_name": "trentb",
      "author_url": "",
      "post_date": "06/26/2014 00:36:52",
      "content": "<p>I spent most of my time on preprocessing to get features that generalized well before trying stacked generalization or a method to handle covariate shift.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 49742,
      "author_name": "wabbit",
      "author_url": "",
      "post_date": "06/27/2014 23:09:49",
      "content": "<p>Trent,</p>\n<p>Thanks for the confirmation.</p>\n<p>I think I understand what is going on-my intuition is that SG will yield maximum benefit in case of minimally&nbsp;processed signal. I am using the fourier space as basis for pre processing, which I believe is destroying ability to discriminate the subjects. Fourier is extremely&nbsp;sub-optimal for random signal. The feature it generates yields SG useless. My feeling is it will impact covariate shift as well. No data to substantiate, will update if i confirm&nbsp;anything.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 50059,
      "author_name": "rforgione",
      "author_url": "",
      "post_date": "07/05/2014 06:25:48",
      "content": "<p>I'm having trouble visualizing the flow of data from one level to the next in ensemble learning/stack generalization. Admittedly, this is the first time I'm working with this concept, so please bear with me as I haven't fully wrapped my head around it.&nbsp;</p>\n\n<p>I spent some time today going through DH Wolpert's (1992) paper on stacked generalization, and reading through Olivetti et. al. (2014) which is posted on the main page of this competition. I start to get tripped up when I read through the following portion on page 6 of Olivetti's paper:</p>\n\n<ol>\n<li>Train a set of classi&#64257;ers on (portions of) the train data. These classi&#64257;ers&nbsp;are called &#64257;rst-level classi&#64257;ers.</li>\n<li>Collect the output of each classi&#64257;er on each trial of the train and of the&nbsp;test data. These outputs are called &#64257;rst level predictions.</li>\n<li><strong>Create a new second-level dataset with the vector of &#64257;rst-level predictions&nbsp;</strong><strong>for each trial. Care has to be taken so that the predicted value of a&nbsp;</strong><strong>given trial comes from classi&#64257;ers which were not trained on that trial, e.g.&nbsp;</strong><strong>through cross-validation.</strong></li>\n<li>The class-labels of the second-level dataset are the same as the initial&nbsp;dataset.</li>\n<li>A second-level classi&#64257;er is trained on the portion of the second-level dataset&nbsp;related to the train subjects in order to learn how to combine the &#64257;rst-level&nbsp;predictions.</li>\n<li>The second level classi&#64257;er is used to predict the class-labels of the test&nbsp;data as represented in the second-level dataset.</li>\n</ol>\n<p>Based on my interpretation, I would&nbsp;expect the second-level dataset to contain the same number of examples as the original train data (we'll call it m), and one feature for each classifier trained on each one&nbsp;of the 16 subjects. Then again, that would seem to betray the requirement that &quot;[C]are has to be taken so that the predicted value of a&nbsp;given trial comes from classi&#64257;ers which were not trained on that trial, e.g.&nbsp;through cross-validation.&quot; For instance, classifier #1 would have been built using all of the trials associated with subject #1. Not sure how to get around this issue.</p>\n\n<p>I've searched around for every paper I could find today on ensemble learning and stacked generalization and have not been able to figure this out (looked at the two listed above plus Sigletos et. al. (2005) and Ting et. al (1999)). If anyone has any advice it would be very helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 50060,
      "author_name": "wabbit",
      "author_url": "",
      "post_date": "07/05/2014 07:23:39",
      "content": "<p>I had the same confusion. My interpretation&nbsp;is:</p>\n<p>1. For each first level, subject specific classifier k, k=1..16 do the following:</p>\n<p>a .&nbsp;Divide &nbsp;data for each subject, k, into J data sets</p>\n<p>b. create J CV classifier.</p>\n<p>c. Train jth classifier using the rest j-1 data sets.</p>\n<p>d.Create L1 meta data, i.e. 'probabilities' for the data set 'j'&nbsp;&nbsp;using the jth classifier. This resolves the issue you mention</p>\n<p>e. Create a 'master' classifier for the kth subject. This classifier is trained using all the subject specific data.</p>\n<p>f. Use the L1 meta-data in step d to&nbsp;as training set for a L1 classifier</p>\n\n<p>2. For a new test data, run it through the master classifiers for each subject.</p>\n<p>3. Use the probalilities as meta-data to classify using the L1 classifier.</p>\n\n<p>I see a marginal improvement with this.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 50064,
      "author_name": "rforgione",
      "author_url": "",
      "post_date": "07/05/2014 15:06:32",
      "content": "<p>So let me make sure that I understand. Let's say that we have a data set of 3 subjects each with 3 trials for a total of 9 trials.&nbsp;</p>\n\n<p>First, we would focus specifically on subject #1. Within subject 1, we would first focus on trial #1. &nbsp;We would use the other two trials for subject #1 (ie trials 2 and 3) &nbsp;to create a classifier, and then we would use this classifier to predict trial #1.&nbsp;</p>\n\n<p>Then, we would move on to trial #2, where we would use trials 1 and 3 within subject #1 to make a classifier, and then use that classifier to predict #2.&nbsp;</p>\n\n<p>We we would repeat this process one more time for trial #3., using 1 and 2 to build the classifier and making a prediction for trial #3.</p>\n\n<p>At this point, we are done with subject #1 and should have 3 predictions, one for each of subject 1's trials.</p>\n\n<p>We would then follow this process for subjects 2 and 3. &nbsp;Ultimately, we end up with a vector of m &nbsp;= 9 predictions, each of which used the other trials within the same subject to build the classifier.</p>\n\n<p>Then we could use this vector to train a new classifier, using the predictions from the preceding step as inputs, and using the original classes as outputs. &nbsp;</p>\n\n<p>Is this correct, or did I misunderstand?</p>\n\n<p>Thanks very much for your help!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 50065,
      "author_name": "rforgione",
      "author_url": "",
      "post_date": "07/05/2014 15:14:41",
      "content": "<p>Although I'm not sure I understand the piece on building a 'master classifier' for each subject..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 50071,
      "author_name": "wabbit",
      "author_url": "",
      "post_date": "07/06/2014 05:48:22",
      "content": "<p>Rob,</p>\n<p>Correct understanding. What classifier model would you use to generate L1 meta data on a new test data? You need one single model. This is the master classifier.&nbsp;</p>\n<p>Disclaimer-this is my understanding and i do not have any concrete data to prove that this is correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 51270,
      "author_name": "choudharydhruv",
      "author_url": "",
      "post_date": "07/28/2014 17:34:00",
      "content": "<p>I implemented stack generalization with some mixed results. Although our final result was better with a pooled model. Here is what I did,</p>\n\n<p>1) Each subject had about 600 trials, I split that into sets of 500 and 100. I trained a classifier on each of the subjects using just the 500 trials.&nbsp;So now I had 16 classifiers. </p>\n<p>2) I used the remaining 100 trials from each to make predictions on each of the 16 classifiers. So this gives me 100*16 trials for my second level classifier. </p>\n<p>3) I ran the whole test set through each of the first level classifiers, used the predictions on the second level classifier.</p>\n<p>This gave me mixed results. Some subjects bumped in accuracy by as much as 0.63-0.74</p>\n<p>Certain subjects like subject 3,16 were pretty hard to predict..they went worse. One the leaderboard this scheme was giving me just about 0.67...so we abandoned it but this was definitely worth trying more.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "49504": "",
    "49520": "",
    "49534": "",
    "49595": "",
    "49742": "",
    "50059": "",
    "50060": "",
    "50064": "",
    "50065": "",
    "50071": "",
    "51270": ""
  },
  "source": "meta"
}