{
  "id": 7919,
  "title": "Useful transformations",
  "url": "/competitions/decoding-the-human-brain/discussion/7919",
  "author_name": "",
  "post_date": "2014-04-29T09:17:45.283Z",
  "votes": null,
  "comment_count": 6,
  "views": 1933,
  "content": "<p>Here you can discuss about useful transformation on MEG data to extract higher level patterns.</p>",
  "messages": [
    {
      "id": "43279",
      "postDate": "04/29/2014 09:17:45",
      "content": "<p>Here you can discuss about useful transformation on MEG data to extract higher level patterns.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43352",
      "postDate": "04/30/2014 09:39:50",
      "content": "<p>Are we allowed to do source reconstruction? Presumably it would involve using the full sensor layout, with 3D positions, which can be easily found in Fieldtrip. So would this violate the terms of the competition?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43427",
      "postDate": "05/01/2014 06:17:16",
      "content": "<p>Hi,</p>\n<p>We are discussing this point. An answer will follow soon.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43442",
      "postDate": "05/01/2014 12:43:24",
      "content": "<p>Hi! I would assume that, in the spirit of good data competitions, any transformation of the data is allowed, right? All that should count&nbsp;is the final predicted result :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43514",
      "postDate": "05/02/2014 14:42:33",
      "content": "<p>Dear All,</p>\n<p>After some discussion among the people organizing this competition, we decided to <strong>accept&nbsp;all possible transformations that the participants will propose</strong>, including source reconstruction, i.e. approximating the currents/dipoles of the brain that generated the observed magnetic field measured by the MEG.</p>\n<p>A potential problem with source reconstruction is the use of external information not provided in the data page of this competition - which is forbidden by the rules. For example: the detailed sensor layout available from some MEG data analysis libraries, or even more detailed&nbsp;information (available <strong>only</strong> for the training subjects) from&nbsp;the original dataset from which the data of this competition is extracted. Since it is difficult to define a clear boundary here, we decided to accept everything. In some sense, this decision weakens the rule of not using external information.</p>\n<p>We kindly invite the participants expert on source reconstruction to share code with all the other participants. We - competition hosts - do not have plan to release example code for source reconstruction.</p>\n<p>A final note: decoding from the reconstructed sources instead of&nbsp;the MEG data, is a recent and almost&nbsp;unexplored trend in MEG data analysis. Our preliminary practical experience is that there is no clear gain in that direction, at least for the problem of this competition. But again, this is almost unknown territory.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44054",
      "postDate": "05/06/2014 15:36:58",
      "content": "<p>Another question on a possible transformation. It is of course standard practice to transform the input features before classification such that they are expressed in the same unit, e.g. z-transform or a scaling between 0 and 1. Usually in machine learning one&nbsp;would do this transform based *only* on quantities estimated from the training set (i.e. estimate mean and sd of training set, and compute z(Xtest) as z(Xtest) = (Xtest - meanTrain) / sdTrain). However, in this particular competition we have access to all the testing data beforehand, i.e. there is no requirement to classify entirely unseen data. Therefore we could do a z-scoring using all data, test and train combined.</p>\n<p>I believe this should not be allowed, as in effect we are then using test data to transform the training set (it has been shown that this can significantly improve&nbsp;classifier accuracy). Could an admin confirm that this is indeed not allowed? Thanks!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "44152",
      "postDate": "05/07/2014 09:10:05",
      "content": "<p>Hi Eelke,</p>\n<p>I disagree with you. In this competition you have MEG measurements (X) and the associated categories of stimulus (y). As long as you don't use information about &quot;y&quot; of the test set, there is no issue or unwanted bias in the analysis that generates the submission. So feel free to use X of the test set. If you are concerned about circularity / double-dipping in the estimate of the score then you shouldn't.</p>\n<p>This competition is about predicting y on new subjects and it is a known fact that the X of different subjects is systematically different. For this reason, one possible approach to the create an effective prediction is exactly to use the X of the test subject in order to adapt (or to &quot;transfer&quot;, as some literature names it) the knowledge that you get from (X,y) of the train subject(s).</p>\n<p>In a standard classification setting, to which you refer in your message, you want to create a classifier that is not specific of a single test set. For this reason you don't want use X of the test set in the process. But this is valid as long as you assume that both the train and test sets come from the same distribution. And in that case you are interested in the expected score over future examples (here trials). If that is not the case - as it is here - then there are other approaches which might be&nbsp;more effective than the standard one - and they use X of the test set. For example semi-supervised learning and transfer learning do that.</p>\n<p>Notice that the scientific questions that we are addressing with this competition is: &quot;what is the prediction accuracy that we should expect on new subjects?&quot;. So it is fine if participants want to adapt the same method for&nbsp;each subject in the test set. Because we are interested in the expected accuracy over subjects.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 43352,
      "author_name": "mirceastoica",
      "author_url": "",
      "post_date": "04/30/2014 09:39:50",
      "content": "<p>Are we allowed to do source reconstruction? Presumably it would involve using the full sensor layout, with 3D positions, which can be easily found in Fieldtrip. So would this violate the terms of the competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43427,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/01/2014 06:17:16",
      "content": "<p>Hi,</p>\n<p>We are discussing this point. An answer will follow soon.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43442,
      "author_name": "eelkespaak",
      "author_url": "",
      "post_date": "05/01/2014 12:43:24",
      "content": "<p>Hi! I would assume that, in the spirit of good data competitions, any transformation of the data is allowed, right? All that should count&nbsp;is the final predicted result :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43514,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/02/2014 14:42:33",
      "content": "<p>Dear All,</p>\n<p>After some discussion among the people organizing this competition, we decided to <strong>accept&nbsp;all possible transformations that the participants will propose</strong>, including source reconstruction, i.e. approximating the currents/dipoles of the brain that generated the observed magnetic field measured by the MEG.</p>\n<p>A potential problem with source reconstruction is the use of external information not provided in the data page of this competition - which is forbidden by the rules. For example: the detailed sensor layout available from some MEG data analysis libraries, or even more detailed&nbsp;information (available <strong>only</strong> for the training subjects) from&nbsp;the original dataset from which the data of this competition is extracted. Since it is difficult to define a clear boundary here, we decided to accept everything. In some sense, this decision weakens the rule of not using external information.</p>\n<p>We kindly invite the participants expert on source reconstruction to share code with all the other participants. We - competition hosts - do not have plan to release example code for source reconstruction.</p>\n<p>A final note: decoding from the reconstructed sources instead of&nbsp;the MEG data, is a recent and almost&nbsp;unexplored trend in MEG data analysis. Our preliminary practical experience is that there is no clear gain in that direction, at least for the problem of this competition. But again, this is almost unknown territory.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44054,
      "author_name": "eelkespaak",
      "author_url": "",
      "post_date": "05/06/2014 15:36:58",
      "content": "<p>Another question on a possible transformation. It is of course standard practice to transform the input features before classification such that they are expressed in the same unit, e.g. z-transform or a scaling between 0 and 1. Usually in machine learning one&nbsp;would do this transform based *only* on quantities estimated from the training set (i.e. estimate mean and sd of training set, and compute z(Xtest) as z(Xtest) = (Xtest - meanTrain) / sdTrain). However, in this particular competition we have access to all the testing data beforehand, i.e. there is no requirement to classify entirely unseen data. Therefore we could do a z-scoring using all data, test and train combined.</p>\n<p>I believe this should not be allowed, as in effect we are then using test data to transform the training set (it has been shown that this can significantly improve&nbsp;classifier accuracy). Could an admin confirm that this is indeed not allowed? Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 44152,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "05/07/2014 09:10:05",
      "content": "<p>Hi Eelke,</p>\n<p>I disagree with you. In this competition you have MEG measurements (X) and the associated categories of stimulus (y). As long as you don't use information about &quot;y&quot; of the test set, there is no issue or unwanted bias in the analysis that generates the submission. So feel free to use X of the test set. If you are concerned about circularity / double-dipping in the estimate of the score then you shouldn't.</p>\n<p>This competition is about predicting y on new subjects and it is a known fact that the X of different subjects is systematically different. For this reason, one possible approach to the create an effective prediction is exactly to use the X of the test subject in order to adapt (or to &quot;transfer&quot;, as some literature names it) the knowledge that you get from (X,y) of the train subject(s).</p>\n<p>In a standard classification setting, to which you refer in your message, you want to create a classifier that is not specific of a single test set. For this reason you don't want use X of the test set in the process. But this is valid as long as you assume that both the train and test sets come from the same distribution. And in that case you are interested in the expected score over future examples (here trials). If that is not the case - as it is here - then there are other approaches which might be&nbsp;more effective than the standard one - and they use X of the test set. For example semi-supervised learning and transfer learning do that.</p>\n<p>Notice that the scientific questions that we are addressing with this competition is: &quot;what is the prediction accuracy that we should expect on new subjects?&quot;. So it is fine if participants want to adapt the same method for&nbsp;each subject in the test set. Because we are interested in the expected accuracy over subjects.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "43279": "",
    "43352": "",
    "43427": "",
    "43442": "",
    "43514": "",
    "44054": "",
    "44152": ""
  },
  "source": "meta"
}