{
  "id": 4468,
  "title": "what does this data mean ?",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4468",
  "author_name": "",
  "post_date": "2013-04-29T16:55:32.217Z",
  "votes": null,
  "comment_count": 11,
  "views": 3154,
  "content": "<p>Of course we'll know what this problem really represents once the competition is over, but it's quite fun to imagine what it could mean before :)</p>\r\n<p>Here is my (wild) guess :</p>\r\n<p>Suppose the classes are cutoffs of a learning algorithm (like their label could imply). Imagine the 1875 features are the weights/outputs of a neural network (as the theme of the ICML workshop : &quot;relational learning&quot; might imply). So the problem would be\r\n to learn to classify the outcome of a (potentially complex) neural structure given its activation weights ! If one competitor succeeds with a comfortable accuracy, this would mean that the structure itself can be replaced by a (potentially totally different)\r\n learner...</p>\r\n<p>Of course my guess might be totally delirious, but hell I thought the idea to be quite funny :)</p>\r\n<p>And you what do you think this data could mean ?</p>",
  "messages": [
    {
      "id": "23642",
      "postDate": "04/29/2013 16:55:32",
      "content": "<p>Of course we'll know what this problem really represents once the competition is over, but it's quite fun to imagine what it could mean before :)</p>\r\n<p>Here is my (wild) guess :</p>\r\n<p>Suppose the classes are cutoffs of a learning algorithm (like their label could imply). Imagine the 1875 features are the weights/outputs of a neural network (as the theme of the ICML workshop : &quot;relational learning&quot; might imply). So the problem would be\r\n to learn to classify the outcome of a (potentially complex) neural structure given its activation weights ! If one competitor succeeds with a comfortable accuracy, this would mean that the structure itself can be replaced by a (potentially totally different)\r\n learner...</p>\r\n<p>Of course my guess might be totally delirious, but hell I thought the idea to be quite funny :)</p>\r\n<p>And you what do you think this data could mean ?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23644",
      "postDate": "04/29/2013 16:58:08",
      "content": "<p>For the record, the title of this workshop isn't &quot;relational learning,&quot; it's &quot;representation learning.&quot;&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23646",
      "postDate": "04/29/2013 17:02:30",
      "content": "<p>sorry, my fingers criss-crossed</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23659",
      "postDate": "04/29/2013 21:31:17",
      "content": "<p>this data is the conversion from image to pixels and then scaling of the inputs.</p>\r\n<p>By the way, i won't be surprised if this data and the face recognition one are the same. With just some image rescaling. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23660",
      "postDate": "04/29/2013 21:35:55",
      "content": "<p>We weren't that clever! :) The (unobfuscated) tasks are quite different.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23663",
      "postDate": "04/29/2013 21:45:24",
      "content": "<p>You guys should have thought this way. So you guys would have menas to compare a blackbox approach and a white box approach on the same problem.&nbsp;<span style=\"line-height:1.4em\">Just needed to select disjoints sets for each competition.&nbsp;</span></p>\r\n<p>But i still believe its a image processing task. No wonder neural nets works best. That and the fact that the labels probably were built with NNs and pylearn2 (hence the bias in the scores!)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23667",
      "postDate": "04/29/2013 21:53:57",
      "content": "<p>We didn't use NNs or pylearn2 to build the labels. The task is designed not to be biased in favor of any machine learning algorithm or another, beyond the choice of the number of labels possibly making it advantageous to use semisupervised learning.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23668",
      "postDate": "04/29/2013 21:57:22",
      "content": "<p>[quote=Leustagos;23663]</p>\r\n<p><span style=\"line-height:1.4em\">No wonder neural nets works best.&nbsp;</span></p>\r\n<p>[/quote]</p>\r\n<p>I haven't really seen the top 3 teams expose their methods here on the forum, so I'm not really sure which method works best for now, to be honest!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23777",
      "postDate": "05/01/2013 21:57:53",
      "content": "<p>I think that the data has been largely artificially generated or modified, I think this because the correlation matrix is realy quite strange, the maximum absolute correlation for some attributes, relative to the other ones, is very high while it is very\r\n low for others and there is every step in between. feels like some artificially added noise with different intensities.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23799",
      "postDate": "05/02/2013 08:50:10",
      "content": "<p>I reckon it's image data</p>\r\n<p>&nbsp;</p>\r\n<p>EDIT: and to be really specific I started thinking too much on it and thought could be 75*25 Facebook cover photos!!!\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23802",
      "postDate": "05/02/2013 12:07:19",
      "content": "<p>[quote=Domcastro;23799]</p>\r\n<p>I reckon it's image data</p>\r\n<p>[/quote]</p>\r\n<p>I was thinking it might be image data, too, possibly 25x25 RGB or HSV images (because both of those colorspaces use triples for each pixel).&nbsp; In any case ithe 25x25x3 (or 5x5x5x5x3) nature of 1875 could certainly be a clue... or not... :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23805",
      "postDate": "05/02/2013 14:19:30",
      "content": "<p>My plan for attacking this problem went through all the approaches to obfuscating the &quot;Original&quot; data that I could imagine. &nbsp;I'm particularly interested in how poorly nearest neighbor methods do, and that among the training data a given element shares the\r\n same class with its nearest neighbor less than 30% of the time. &nbsp;A problem begging for semi-supervised learning at first glance. &nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 23644,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 16:58:08",
      "content": "<p>For the record, the title of this workshop isn't &quot;relational learning,&quot; it's &quot;representation learning.&quot;&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23646,
      "author_name": "eustache",
      "author_url": "",
      "post_date": "04/29/2013 17:02:30",
      "content": "<p>sorry, my fingers criss-crossed</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23659,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/29/2013 21:31:17",
      "content": "<p>this data is the conversion from image to pixels and then scaling of the inputs.</p>\r\n<p>By the way, i won't be surprised if this data and the face recognition one are the same. With just some image rescaling. &nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23660,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "04/29/2013 21:35:55",
      "content": "<p>We weren't that clever! :) The (unobfuscated) tasks are quite different.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23663,
      "author_name": "leustagos",
      "author_url": "",
      "post_date": "04/29/2013 21:45:24",
      "content": "<p>You guys should have thought this way. So you guys would have menas to compare a blackbox approach and a white box approach on the same problem.&nbsp;<span style=\"line-height:1.4em\">Just needed to select disjoints sets for each competition.&nbsp;</span></p>\r\n<p>But i still believe its a image processing task. No wonder neural nets works best. That and the fact that the labels probably were built with NNs and pylearn2 (hence the bias in the scores!)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23667,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/29/2013 21:53:57",
      "content": "<p>We didn't use NNs or pylearn2 to build the labels. The task is designed not to be biased in favor of any machine learning algorithm or another, beyond the choice of the number of labels possibly making it advantageous to use semisupervised learning.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23668,
      "author_name": "dumitru0",
      "author_url": "",
      "post_date": "04/29/2013 21:57:22",
      "content": "<p>[quote=Leustagos;23663]</p>\r\n<p><span style=\"line-height:1.4em\">No wonder neural nets works best.&nbsp;</span></p>\r\n<p>[/quote]</p>\r\n<p>I haven't really seen the top 3 teams expose their methods here on the forum, so I'm not really sure which method works best for now, to be honest!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23777,
      "author_name": "emericbeaufays",
      "author_url": "",
      "post_date": "05/01/2013 21:57:53",
      "content": "<p>I think that the data has been largely artificially generated or modified, I think this because the correlation matrix is realy quite strange, the maximum absolute correlation for some attributes, relative to the other ones, is very high while it is very\r\n low for others and there is every step in between. feels like some artificially added noise with different intensities.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23799,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "05/02/2013 08:50:10",
      "content": "<p>I reckon it's image data</p>\r\n<p>&nbsp;</p>\r\n<p>EDIT: and to be really specific I started thinking too much on it and thought could be 75*25 Facebook cover photos!!!\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23802,
      "author_name": "yetiman",
      "author_url": "",
      "post_date": "05/02/2013 12:07:19",
      "content": "<p>[quote=Domcastro;23799]</p>\r\n<p>I reckon it's image data</p>\r\n<p>[/quote]</p>\r\n<p>I was thinking it might be image data, too, possibly 25x25 RGB or HSV images (because both of those colorspaces use triples for each pixel).&nbsp; In any case ithe 25x25x3 (or 5x5x5x5x3) nature of 1875 could certainly be a clue... or not... :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23805,
      "author_name": "",
      "author_url": "",
      "post_date": "05/02/2013 14:19:30",
      "content": "<p>My plan for attacking this problem went through all the approaches to obfuscating the &quot;Original&quot; data that I could imagine. &nbsp;I'm particularly interested in how poorly nearest neighbor methods do, and that among the training data a given element shares the\r\n same class with its nearest neighbor less than 30% of the time. &nbsp;A problem begging for semi-supervised learning at first glance. &nbsp;</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "23642": "",
    "23644": "",
    "23646": "",
    "23659": "",
    "23660": "",
    "23663": "",
    "23667": "",
    "23668": "",
    "23777": "",
    "23799": "",
    "23802": "",
    "23805": ""
  },
  "source": "meta"
}