{
  "id": 4319,
  "title": "Lack of training data",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4319",
  "author_name": "",
  "post_date": "2013-04-15T13:45:24.140Z",
  "votes": null,
  "comment_count": 6,
  "views": 1997,
  "content": "<p>The challenge for this competition isn't the lack of information about the data but just the lack of training data.&nbsp; Not sure why it's called the &quot;black box learning challenge&quot; - would be more deserved to be called &quot;Lack of training data&quot; challenge.</p>\r\n<p>&nbsp;</p>\r\n<p>The winners would have proved that they are the best at working with minimal training data. Lack of data knowledge seems to be irrelevant (which defeats the whole object of the competition)</p>",
  "messages": [
    {
      "id": "22834",
      "postDate": "04/15/2013 13:45:24",
      "content": "<p>The challenge for this competition isn't the lack of information about the data but just the lack of training data.&nbsp; Not sure why it's called the &quot;black box learning challenge&quot; - would be more deserved to be called &quot;Lack of training data&quot; challenge.</p>\r\n<p>&nbsp;</p>\r\n<p>The winners would have proved that they are the best at working with minimal training data. Lack of data knowledge seems to be irrelevant (which defeats the whole object of the competition)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22835",
      "postDate": "04/15/2013 14:07:48",
      "content": "<p>You have plenty of training data--over 100,000 examples. What you don't have a lot of is *labeled* training data.</p>\r\n<p>In a competition with few labels and a known data domain, a human beings would be able to design good features for a problem with few labels, based on their knowledge of the data domain. Here, you don't have knowledge of the data domain, so it's up to your\r\n algorithm to build that knowledge from the unlabeled data.</p>\r\n<p>With the right features for a task, it's entirely possible to generalize well to new data using only a few examples. Consider for example the popular Iris dataset. It's based on classifying species of flowers from measurements of parts of the plant. These\r\n measurements are a good feature space in which to classify the plant so even a simple model like softmax regression can learn to classify essentially perfectly, with fewer than 150 training cases.\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22842",
      "postDate": "04/15/2013 15:06:27",
      "content": "<p>I don't consider unlabelled data as training data. Most of the competitions on Kaggle don't require an expert so I'm not sure why this one is any different. (apart from the lack of training data) So the challenge is semi-supervised learning rather than black\r\n box</p>\r\n<p>EDIT: Sorry - just was excited about not having to spend intellectual effort on this competition (ie just data mining) but now realise I have to spend more intellectual effort learning semi-supervised methods. No time for this effort :(\r\n</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22848",
      "postDate": "04/15/2013 16:40:11",
      "content": "<p>Yes, this contest is for a scientific conference, so some intellectual effort is do expected. Specifically, it's for a workshop on representation learning, so we do expect semi-supervised methods to perform well.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22850",
      "postDate": "04/15/2013 17:02:41",
      "content": "<p>I'd just like to add that, though I'm not sure I'll have time to compete in this competition, I'm very glad to see something covering un/semi-supervised learning on Kaggle as I find that problems with a high unlabeld/labeld data ratio are extremely common\r\n in the real world. &nbsp;This is especially true for problems that require human expertise for labeling as human experts tend to take a long time to label data and are also very costly. &nbsp;Not only is this an interesting research question, it's also extremely applicable\r\n to a wide range of intesting data challenges out there.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23753",
      "postDate": "05/01/2013 15:45:15",
      "content": "<p>Class &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;count</p>\r\n<table border=\"0\" cellspacing=\"0\" cellpadding=\"0\" style=\"width:162px\">\r\n<colgroup><col width=\"92\"><col width=\"70\"></colgroup>\r\n<tbody>\r\n<tr>\r\n<td class=\"x_xl63\" width=\"92\" height=\"20\">9</td>\r\n<td align=\"right\" width=\"70\">75</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">8</td>\r\n<td align=\"right\">73</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">7</td>\r\n<td align=\"right\">81</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">6</td>\r\n<td align=\"right\">82</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">5</td>\r\n<td align=\"right\">80</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">4</td>\r\n<td align=\"right\">111</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">3</td>\r\n<td align=\"right\">130</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">2</td>\r\n<td align=\"right\">170</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">1</td>\r\n<td align=\"right\">197</td>\r\n</tr>\r\n</tbody>\r\n</table>\r\n<p>Wish the count of objects in each class were consistent.&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "23754",
      "postDate": "05/01/2013 15:48:33",
      "content": "<p>The real world isn't always sorted into nicely balanced classes.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 22835,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/15/2013 14:07:48",
      "content": "<p>You have plenty of training data--over 100,000 examples. What you don't have a lot of is *labeled* training data.</p>\r\n<p>In a competition with few labels and a known data domain, a human beings would be able to design good features for a problem with few labels, based on their knowledge of the data domain. Here, you don't have knowledge of the data domain, so it's up to your\r\n algorithm to build that knowledge from the unlabeled data.</p>\r\n<p>With the right features for a task, it's entirely possible to generalize well to new data using only a few examples. Consider for example the popular Iris dataset. It's based on classifying species of flowers from measurements of parts of the plant. These\r\n measurements are a good feature space in which to classify the plant so even a simple model like softmax regression can learn to classify essentially perfectly, with fewer than 150 training cases.\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22842,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "04/15/2013 15:06:27",
      "content": "<p>I don't consider unlabelled data as training data. Most of the competitions on Kaggle don't require an expert so I'm not sure why this one is any different. (apart from the lack of training data) So the challenge is semi-supervised learning rather than black\r\n box</p>\r\n<p>EDIT: Sorry - just was excited about not having to spend intellectual effort on this competition (ie just data mining) but now realise I have to spend more intellectual effort learning semi-supervised methods. No time for this effort :(\r\n</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22848,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "04/15/2013 16:40:11",
      "content": "<p>Yes, this contest is for a scientific conference, so some intellectual effort is do expected. Specifically, it's for a workshop on representation learning, so we do expect semi-supervised methods to perform well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22850,
      "author_name": "willkurt",
      "author_url": "",
      "post_date": "04/15/2013 17:02:41",
      "content": "<p>I'd just like to add that, though I'm not sure I'll have time to compete in this competition, I'm very glad to see something covering un/semi-supervised learning on Kaggle as I find that problems with a high unlabeld/labeld data ratio are extremely common\r\n in the real world. &nbsp;This is especially true for problems that require human expertise for labeling as human experts tend to take a long time to label data and are also very costly. &nbsp;Not only is this an interesting research question, it's also extremely applicable\r\n to a wide range of intesting data challenges out there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23753,
      "author_name": "kosmos",
      "author_url": "",
      "post_date": "05/01/2013 15:45:15",
      "content": "<p>Class &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;count</p>\r\n<table border=\"0\" cellspacing=\"0\" cellpadding=\"0\" style=\"width:162px\">\r\n<colgroup><col width=\"92\"><col width=\"70\"></colgroup>\r\n<tbody>\r\n<tr>\r\n<td class=\"x_xl63\" width=\"92\" height=\"20\">9</td>\r\n<td align=\"right\" width=\"70\">75</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">8</td>\r\n<td align=\"right\">73</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">7</td>\r\n<td align=\"right\">81</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">6</td>\r\n<td align=\"right\">82</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">5</td>\r\n<td align=\"right\">80</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">4</td>\r\n<td align=\"right\">111</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">3</td>\r\n<td align=\"right\">130</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">2</td>\r\n<td align=\"right\">170</td>\r\n</tr>\r\n<tr>\r\n<td class=\"x_xl63\" height=\"20\">1</td>\r\n<td align=\"right\">197</td>\r\n</tr>\r\n</tbody>\r\n</table>\r\n<p>Wish the count of objects in each class were consistent.&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 23754,
      "author_name": "iangoodfellow",
      "author_url": "",
      "post_date": "05/01/2013 15:48:33",
      "content": "<p>The real world isn't always sorted into nicely balanced classes.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "22834": "",
    "22835": "",
    "22842": "",
    "22848": "",
    "22850": "",
    "23753": "",
    "23754": ""
  },
  "source": "meta"
}