{
  "id": 8243,
  "title": "Covariate shift",
  "url": "/competitions/decoding-the-human-brain/discussion/8243",
  "author_name": "",
  "post_date": "2014-05-21T12:03:50.880Z",
  "votes": null,
  "comment_count": 4,
  "views": 2065,
  "content": "<p>Hi,</p>\n<p>can someone help me understand how to implement the covariate shift described in the paper?</p>\n<p>I get that the idea is to train a classifier to distinguish between test and train sets and weight each instance with p(test)/p(train). I am using liblinear (of the makers of libsvm) to output probability estimates with&nbsp;logistic regression.&nbsp;</p>\n<p>Please correct any steps I got wrong&nbsp;in leave one subject out cross-validation:&nbsp;</p>\n<p>For each subject:</p>\n<p>1. Train set is all the other subjects, test set is the current subject.</p>\n<p>2. Train the covariate shift classifier and output probability estimates for each instance through cross-validation. So we take K-1 folds of both train and test sets from 1 and predict the probability estimates for the remaining fold. Regardless of the stimulus condition, the whole train set from 1 will have the same label and the whole test set will have a different label.</p>\n<p>3. Multiply each instance from 1 with p(test)/p(train) obtained from 2.</p>\n<p>4. Train the standard classifier to distinguish the stimuli on the weighted train set&nbsp;from 3 and predict on the weighted test set.</p>\n<p>This approach only led to lower CV accuracies. I've also tried setting different costs for the two classes in 2, as good practice calls for when using SVMs, but still no joy.</p>\n<p>It would be much appreciated and probably helpful for others as well if someone would clear this up.</p>",
  "messages": [
    {
      "id": "45983",
      "postDate": "05/21/2014 12:03:50",
      "content": "<p>Hi,</p>\n<p>can someone help me understand how to implement the covariate shift described in the paper?</p>\n<p>I get that the idea is to train a classifier to distinguish between test and train sets and weight each instance with p(test)/p(train). I am using liblinear (of the makers of libsvm) to output probability estimates with&nbsp;logistic regression.&nbsp;</p>\n<p>Please correct any steps I got wrong&nbsp;in leave one subject out cross-validation:&nbsp;</p>\n<p>For each subject:</p>\n<p>1. Train set is all the other subjects, test set is the current subject.</p>\n<p>2. Train the covariate shift classifier and output probability estimates for each instance through cross-validation. So we take K-1 folds of both train and test sets from 1 and predict the probability estimates for the remaining fold. Regardless of the stimulus condition, the whole train set from 1 will have the same label and the whole test set will have a different label.</p>\n<p>3. Multiply each instance from 1 with p(test)/p(train) obtained from 2.</p>\n<p>4. Train the standard classifier to distinguish the stimuli on the weighted train set&nbsp;from 3 and predict on the weighted test set.</p>\n<p>This approach only led to lower CV accuracies. I've also tried setting different costs for the two classes in 2, as good practice calls for when using SVMs, but still no joy.</p>\n<p>It would be much appreciated and probably helpful for others as well if someone would clear this up.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48628",
      "postDate": "06/04/2014 14:02:53",
      "content": "<p>Emanuele, since it is your paper, would you please shed some light on the matter? If it works it is of undebatable importance to everyone, yet the precise way of weighting training instances is unclear.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48701",
      "postDate": "06/05/2014 10:40:55",
      "content": "<p>Hi,</p>\n<p>The paper describes the problem of decoding across subjects as a transfer learning / adaptation problem. As a first attempt to implement that idea, a combination of stacked generalisation and simple covariate shift is proposed. Stacked generalisation helps in two ways: it models (somewhat) the variability across train subjects and it greatly reduces the dimension of the feature space (from thousands to tens). This second aspect may be important for simple covariate shift, in my opinion.</p>\n<p>In your description you report&nbsp;only the simple covariate shift part. In my experiments, using only simple covariate shift on the original feature space did not work at all. My intuition is that estimating p(test) or p(train) is more difficult in high dimensions. Moreover, consider that the imbalancedness of the dataset may cause problems.</p>\n<p>As a final comment, I consider the proposed algorithm&nbsp;of the paper as a very preliminary solution. It is meant to show a basic attempt in the topic of transfer learning on MEG data. I believe that its score in this competition would be far from top positions&nbsp;and that a large part of the participants are already doing better than that. In other words, the main message of the paper is &quot;this problem falls in the transfer learning domain&quot;, more than proposing a very effective algorithm.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48707",
      "postDate": "06/05/2014 11:53:33",
      "content": "<p>Thanks Emanuele!</p>\n\n<p>My attempts at covariate shift correction are based on the second level dataset, after stacked generalization (which does improve performance).</p>\n\n<p>What I still don't understand is how to weight instances...does one simply multiply them by p(test)/p(train)? Should both train and test sets be weighted or only the train set? Whatever I tried (on the second level dataset, after stacked generalization) only led to lower classification accuracy so I can't figure out empirically which is the correct way to do it.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "48711",
      "postDate": "06/05/2014 12:46:38",
      "content": "<p>Hi,</p>\n<p>As proposed in many articles in the literature, the empirical risk minimisation framework provides a way to assign weights to training examples. This means that during training you can minimise the weighted loss of each example instead of the stadard&nbsp;loss (as reported in Eq.1-2-3 in the <a href=\"http://arxiv.org/abs/1404.4175\">paper</a>). In order put this in practice, either you rely on a software library that allows it, e.g. <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.SGDClassifier.html#sklearn.linear_model.SGDClassifier.fit\">scikit-learn &quot;sample_weight&quot; parameter</a>&nbsp;(available for some classifiers), or you implement the minimisation function yourself with what you need.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 48628,
      "author_name": "mirceastoica",
      "author_url": "",
      "post_date": "06/04/2014 14:02:53",
      "content": "<p>Emanuele, since it is your paper, would you please shed some light on the matter? If it works it is of undebatable importance to everyone, yet the precise way of weighting training instances is unclear.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48701,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "06/05/2014 10:40:55",
      "content": "<p>Hi,</p>\n<p>The paper describes the problem of decoding across subjects as a transfer learning / adaptation problem. As a first attempt to implement that idea, a combination of stacked generalisation and simple covariate shift is proposed. Stacked generalisation helps in two ways: it models (somewhat) the variability across train subjects and it greatly reduces the dimension of the feature space (from thousands to tens). This second aspect may be important for simple covariate shift, in my opinion.</p>\n<p>In your description you report&nbsp;only the simple covariate shift part. In my experiments, using only simple covariate shift on the original feature space did not work at all. My intuition is that estimating p(test) or p(train) is more difficult in high dimensions. Moreover, consider that the imbalancedness of the dataset may cause problems.</p>\n<p>As a final comment, I consider the proposed algorithm&nbsp;of the paper as a very preliminary solution. It is meant to show a basic attempt in the topic of transfer learning on MEG data. I believe that its score in this competition would be far from top positions&nbsp;and that a large part of the participants are already doing better than that. In other words, the main message of the paper is &quot;this problem falls in the transfer learning domain&quot;, more than proposing a very effective algorithm.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48707,
      "author_name": "mirceastoica",
      "author_url": "",
      "post_date": "06/05/2014 11:53:33",
      "content": "<p>Thanks Emanuele!</p>\n\n<p>My attempts at covariate shift correction are based on the second level dataset, after stacked generalization (which does improve performance).</p>\n\n<p>What I still don't understand is how to weight instances...does one simply multiply them by p(test)/p(train)? Should both train and test sets be weighted or only the train set? Whatever I tried (on the second level dataset, after stacked generalization) only led to lower classification accuracy so I can't figure out empirically which is the correct way to do it.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 48711,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "06/05/2014 12:46:38",
      "content": "<p>Hi,</p>\n<p>As proposed in many articles in the literature, the empirical risk minimisation framework provides a way to assign weights to training examples. This means that during training you can minimise the weighted loss of each example instead of the stadard&nbsp;loss (as reported in Eq.1-2-3 in the <a href=\"http://arxiv.org/abs/1404.4175\">paper</a>). In order put this in practice, either you rely on a software library that allows it, e.g. <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.linear_model.SGDClassifier.html#sklearn.linear_model.SGDClassifier.fit\">scikit-learn &quot;sample_weight&quot; parameter</a>&nbsp;(available for some classifiers), or you implement the minimisation function yourself with what you need.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "45983": "",
    "48628": "",
    "48701": "",
    "48707": "",
    "48711": ""
  },
  "source": "meta"
}