{
  "id": 7920,
  "title": "preprocessing details - low rank of data?",
  "url": "/competitions/decoding-the-human-brain/discussion/7920",
  "author_name": "",
  "post_date": "2014-04-29T09:47:54.747Z",
  "votes": 1,
  "comment_count": 5,
  "views": 1401,
  "content": "<p>I was wondering if it would be possible for the organizers to provide some more details on the preprocessing steps applied to the data.</p>\n<p>I have found the rank of the raw data to be very variable between subjects and much lower than I would expect - which suggests some extensive preprocessing has been done. 5 subjects have data with rank less than 10 (2 of the test subjects) and one has data of rank 2! (subject 9). My understanding is that after the maxfilter procedure rank should be greatly reduced, but probably not to this extent (ie above 30).</p>\n<p>Here are the ranks of the raw competition data for each subject, and below the Matlab code I used to calculate this:</p>\n<p><code><br>Sub: 01 Rank: 201<br>Sub: 02 Rank: 201<br>Sub: 03 Rank: 202<br>Sub: 04 Rank: 33<br>Sub: 05 Rank: 23<br>Sub: 06 Rank: 33<br>Sub: 07 Rank: 20<br>Sub: 08 Rank: 138<br>Sub: 09 Rank: 2<br>Sub: 10 Rank: 8<br>Sub: 11 Rank: 20<br>Sub: 12 Rank: 13<br>Sub: 13 Rank: 146<br>Sub: 14 Rank: 202<br>Sub: 15 Rank: 7<br>Sub: 16 Rank: 202<br>Sub: 17 Rank: 32<br>Sub: 18 Rank: 4<br>Sub: 19 Rank: 38<br>Sub: 20 Rank: 30<br>Sub: 21 Rank: 164<br>Sub: 22 Rank: 151<br>Sub: 23 Rank: 8</code></p>\n<p><code>subj_type = [repmat({'train'},1,16) repmat({'test'},1,7)];<br>Nsub = 23;<br>for subi=1:Nsub<br> fname = sprintf('%s_subject%02d.mat',subj_type{subi},subi);<br> dat = load(fullfile(data_dir,fname));<br> <br> % permute to channels first<br> tmp = permute(dat.X, [2 3 1]);<br> r = rank(tmp(:,:));<br> <br> fprintf(1,'Sub: %02d Rank: %d\\n', subi, r);<br>end</code></p>",
  "messages": [
    {
      "id": "43280",
      "postDate": "04/29/2014 09:47:54",
      "content": "<p>I was wondering if it would be possible for the organizers to provide some more details on the preprocessing steps applied to the data.</p>\n<p>I have found the rank of the raw data to be very variable between subjects and much lower than I would expect - which suggests some extensive preprocessing has been done. 5 subjects have data with rank less than 10 (2 of the test subjects) and one has data of rank 2! (subject 9). My understanding is that after the maxfilter procedure rank should be greatly reduced, but probably not to this extent (ie above 30).</p>\n<p>Here are the ranks of the raw competition data for each subject, and below the Matlab code I used to calculate this:</p>\n<p><code><br>Sub: 01 Rank: 201<br>Sub: 02 Rank: 201<br>Sub: 03 Rank: 202<br>Sub: 04 Rank: 33<br>Sub: 05 Rank: 23<br>Sub: 06 Rank: 33<br>Sub: 07 Rank: 20<br>Sub: 08 Rank: 138<br>Sub: 09 Rank: 2<br>Sub: 10 Rank: 8<br>Sub: 11 Rank: 20<br>Sub: 12 Rank: 13<br>Sub: 13 Rank: 146<br>Sub: 14 Rank: 202<br>Sub: 15 Rank: 7<br>Sub: 16 Rank: 202<br>Sub: 17 Rank: 32<br>Sub: 18 Rank: 4<br>Sub: 19 Rank: 38<br>Sub: 20 Rank: 30<br>Sub: 21 Rank: 164<br>Sub: 22 Rank: 151<br>Sub: 23 Rank: 8</code></p>\n<p><code>subj_type = [repmat({'train'},1,16) repmat({'test'},1,7)];<br>Nsub = 23;<br>for subi=1:Nsub<br> fname = sprintf('%s_subject%02d.mat',subj_type{subi},subi);<br> dat = load(fullfile(data_dir,fname));<br> <br> % permute to channels first<br> tmp = permute(dat.X, [2 3 1]);<br> r = rank(tmp(:,:));<br> <br> fprintf(1,'Sub: %02d Rank: %d\\n', subi, r);<br>end</code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43294",
      "postDate": "04/29/2014 14:13:06",
      "content": "<p>Hi Robin,</p>\n<p>As pointed out in another thread of this forum,&nbsp;we preferred to keep the pre-processing as simple as possible, when preparing the datasets of the competition. As described in the <a href=\"https://www.kaggle.com/c/decoding-the-human-brain/data\">&quot;get the data&quot; page</a>, we&nbsp;high-pass filtered the raw data&nbsp;at 1Hz, and then downsampled to 200Hz before slicing (epoching) the trials in [-0.5. 1.0]sec time windows from when the stimuli started.</p>\n<p>On the one hand a simple pre-processing does not eliminate some of the issues that are in the data. On the other hand, preprocessing is an open-ended problem so there is no ultimate solution to it. Over the last months we&nbsp;tried several different pre-processing choices (maxfilter too) but we did not see any substantial improvement with respect to the one adopted for the competition. Of course this is not a general claim that extensive pre-processing is of little use. I just report the evidence that we collected in our attempts on this dataset.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43295",
      "postDate": "04/29/2014 14:16:39",
      "content": "<p>So the data used for the competition has not been processed with maxfilter?</p>\n<p>Do you have any idea why the rank of the data is so low?</p>\n<p>I would expect unprocessed data or data minimally processed as you describe to have full rank (ie equal to the number of sensors). </p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43296",
      "postDate": "04/29/2014 14:25:50",
      "content": "<p>Can I ask how you performed the downsampling? Presumably there was a low-pass filtering step beforehand to avoid aliasing?</p>\n<p>I absolutely agree with your point about minimal processing and would also chose to start from as close to the raw data as possible. I am just trying to understand why the rank is so low - one subject has only 2 degrees of freedom in the 306-channel data. How can that be possible if the only thing that has been done to the data is a high pass filter and resampling?</p>\n<p>[EDIT:</p>\n<p>I think I found my mistake! It was problems with numerical precision. When I scale up the data before calculating the rank (and calculate instead the rank of the covariance) the values are much higher - although still far from 306 (so could it still be possible that maxfilter was applied to this dataset?)</p>\n<p><code> tmp = 10e20*permute(dat.X, [2 3 1]);<br> tmp = tmp(:,:);<br> r = rank(cov(tmp'));<br></code></p>\n<p><br>Sub: 01 Rank: 202<br>Sub: 02 Rank: 201<br>Sub: 03 Rank: 205<br>Sub: 04 Rank: 201<br>Sub: 05 Rank: 201<br>Sub: 06 Rank: 202<br>Sub: 07 Rank: 201<br>Sub: 08 Rank: 201<br>Sub: 09 Rank: 49<br>Sub: 10 Rank: 201<br>Sub: 11 Rank: 201<br>Sub: 12 Rank: 52<br>Sub: 13 Rank: 202<br>Sub: 14 Rank: 205<br>Sub: 15 Rank: 148<br>Sub: 16 Rank: 204<br>Sub: 17 Rank: 201<br>Sub: 18 Rank: 20<br>Sub: 19 Rank: 201<br>Sub: 20 Rank: 202<br>Sub: 21 Rank: 201<br>Sub: 22 Rank: 201<br>Sub: 23 Rank: 152</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43369",
      "postDate": "04/30/2014 15:54:54",
      "content": "<p>Hi,</p>\n<p>We re-checked the files and the steps that created the dataset of the competition: the maxfilter was not applied during pre-processing.</p>\n<p>Nevertheless your comment about the rank of the data is interesting and we are looking into that.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "43725",
      "postDate": "05/05/2014 15:47:34",
      "content": "<p>FWIW, I am consistently seeing a rank of 298. This seems to hold for all trials and subjects (although I've only verified it in a limited subset). Testing as follows, for e.g. trial 1:</p>\n<p><code>rank(squeeze(X(1,:,:)))</code><code><br></code></p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 43294,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "04/29/2014 14:13:06",
      "content": "<p>Hi Robin,</p>\n<p>As pointed out in another thread of this forum,&nbsp;we preferred to keep the pre-processing as simple as possible, when preparing the datasets of the competition. As described in the <a href=\"https://www.kaggle.com/c/decoding-the-human-brain/data\">&quot;get the data&quot; page</a>, we&nbsp;high-pass filtered the raw data&nbsp;at 1Hz, and then downsampled to 200Hz before slicing (epoching) the trials in [-0.5. 1.0]sec time windows from when the stimuli started.</p>\n<p>On the one hand a simple pre-processing does not eliminate some of the issues that are in the data. On the other hand, preprocessing is an open-ended problem so there is no ultimate solution to it. Over the last months we&nbsp;tried several different pre-processing choices (maxfilter too) but we did not see any substantial improvement with respect to the one adopted for the competition. Of course this is not a general claim that extensive pre-processing is of little use. I just report the evidence that we collected in our attempts on this dataset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43295,
      "author_name": "robince",
      "author_url": "",
      "post_date": "04/29/2014 14:16:39",
      "content": "<p>So the data used for the competition has not been processed with maxfilter?</p>\n<p>Do you have any idea why the rank of the data is so low?</p>\n<p>I would expect unprocessed data or data minimally processed as you describe to have full rank (ie equal to the number of sensors). </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43296,
      "author_name": "robince",
      "author_url": "",
      "post_date": "04/29/2014 14:25:50",
      "content": "<p>Can I ask how you performed the downsampling? Presumably there was a low-pass filtering step beforehand to avoid aliasing?</p>\n<p>I absolutely agree with your point about minimal processing and would also chose to start from as close to the raw data as possible. I am just trying to understand why the rank is so low - one subject has only 2 degrees of freedom in the 306-channel data. How can that be possible if the only thing that has been done to the data is a high pass filter and resampling?</p>\n<p>[EDIT:</p>\n<p>I think I found my mistake! It was problems with numerical precision. When I scale up the data before calculating the rank (and calculate instead the rank of the covariance) the values are much higher - although still far from 306 (so could it still be possible that maxfilter was applied to this dataset?)</p>\n<p><code> tmp = 10e20*permute(dat.X, [2 3 1]);<br> tmp = tmp(:,:);<br> r = rank(cov(tmp'));<br></code></p>\n<p><br>Sub: 01 Rank: 202<br>Sub: 02 Rank: 201<br>Sub: 03 Rank: 205<br>Sub: 04 Rank: 201<br>Sub: 05 Rank: 201<br>Sub: 06 Rank: 202<br>Sub: 07 Rank: 201<br>Sub: 08 Rank: 201<br>Sub: 09 Rank: 49<br>Sub: 10 Rank: 201<br>Sub: 11 Rank: 201<br>Sub: 12 Rank: 52<br>Sub: 13 Rank: 202<br>Sub: 14 Rank: 205<br>Sub: 15 Rank: 148<br>Sub: 16 Rank: 204<br>Sub: 17 Rank: 201<br>Sub: 18 Rank: 20<br>Sub: 19 Rank: 201<br>Sub: 20 Rank: 202<br>Sub: 21 Rank: 201<br>Sub: 22 Rank: 201<br>Sub: 23 Rank: 152</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43369,
      "author_name": "emanuele",
      "author_url": "",
      "post_date": "04/30/2014 15:54:54",
      "content": "<p>Hi,</p>\n<p>We re-checked the files and the steps that created the dataset of the competition: the maxfilter was not applied during pre-processing.</p>\n<p>Nevertheless your comment about the rank of the data is interesting and we are looking into that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 43725,
      "author_name": "eelkespaak",
      "author_url": "",
      "post_date": "05/05/2014 15:47:34",
      "content": "<p>FWIW, I am consistently seeing a rank of 298. This seems to hold for all trials and subjects (although I've only verified it in a limited subset). Testing as follows, for e.g. trial 1:</p>\n<p><code>rank(squeeze(X(1,:,:)))</code><code><br></code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "43280": "",
    "43294": "",
    "43295": "",
    "43296": "",
    "43369": "",
    "43725": ""
  },
  "source": "meta"
}