{
  "id": 4706,
  "title": "why pseudo-labels?",
  "url": "/competitions/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4706",
  "author_name": "",
  "post_date": "2013-05-30T10:31:33.817Z",
  "votes": 1,
  "comment_count": 4,
  "views": 5096,
  "content": "<p>In the models thread, sayit and&nbsp;Gilberto Titericz Junior mentioned using pseudo-labels for extra data. As far as I understand, you'd train a model, probably a neural &nbsp;network, on labeled data and then predict labels for unlabeled data. Then use the extra\r\n data with those predicted pseudo-labels for training. Why would it work?</p>",
  "messages": [
    {
      "id": "24931",
      "postDate": "05/30/2013 10:31:33",
      "content": "<p>In the models thread, sayit and&nbsp;Gilberto Titericz Junior mentioned using pseudo-labels for extra data. As far as I understand, you'd train a model, probably a neural &nbsp;network, on labeled data and then predict labels for unlabeled data. Then use the extra\r\n data with those predicted pseudo-labels for training. Why would it work?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24932",
      "postDate": "05/30/2013 11:05:04",
      "content": "<p>See this paper entitled 'Entropy Regularization':&nbsp;http://www.iro.umontreal.ca/~lisa/publications2/index.php/publications/show/8</p>\r\n<p>It favors a low-density separation between classes, a commonly assumed prior for classification problems in machine learning.</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24961",
      "postDate": "05/30/2013 18:35:50",
      "content": "<p>The effect of supervised training unlabeled data with pseudo-labels is that the (neural) network outputs of unlabeled data are closer to 1 or 0 than training only labeled data.</p>\r\n<p>I think that this is some kind of contractive regularization, I need to study more for detailed theoretical background (including the article above:))</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "24962",
      "postDate": "05/30/2013 19:29:41",
      "content": "<p></p>\r\n<p>[quote=sayit;24961]</p>\r\n<p>The effect of supervised training unlabeled data with pseudo-labels is that the (neural) network outputs of unlabeled data are closer to 1 or 0 than training only labeled data.</p>\r\n<p>I think that this is some kind of contractive regularization, I need to study more for detailed theoretical background (including the article above:))</p>\r\n<p></p>\r\n<p>[/quote]</p>\r\n<p>Intuitively, a network learns to confirm its own suspicions, so it makes sense.</p>\r\n<p>As far as I understand, the key concept here is a cluster assumption, meaning that each class forms a cluster and the clusters are separated by low density regions. The intro in this paper sums it up nicely:</p>\r\n<p>Chapelle and Zien:&nbsp;Semi-Supervised Classification by Low Density Separation</p>\r\n<p>http://www.kyb.mpg.de/publications/pdfs/pdf2899.pdf</p>\r\n<p></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "25016",
      "postDate": "06/01/2013 01:51:39",
      "content": "<p><span>In my experiments, </span><strong>pseudo-labels are re-calculated every weights update<strong>&nbsp;in training with labeled and unlabeled data\r\n<strong><strong>simultaneously</strong></strong>.</strong></strong>&nbsp;If we calculate pseudo-label once after training with only labeled data, pseudo-label might be less accurate because the network is overfitted. After training several initial epochs with only\r\n labeled data, the network should be trained with labeled data and unlabeled data using continuously re-calculated pseudo-label. This scheme improve the generalization performance really.</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 24932,
      "author_name": "yoshuabengio",
      "author_url": "",
      "post_date": "05/30/2013 11:05:04",
      "content": "<p>See this paper entitled 'Entropy Regularization':&nbsp;http://www.iro.umontreal.ca/~lisa/publications2/index.php/publications/show/8</p>\r\n<p>It favors a low-density separation between classes, a commonly assumed prior for classification problems in machine learning.</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24961,
      "author_name": "donghyun",
      "author_url": "",
      "post_date": "05/30/2013 18:35:50",
      "content": "<p>The effect of supervised training unlabeled data with pseudo-labels is that the (neural) network outputs of unlabeled data are closer to 1 or 0 than training only labeled data.</p>\r\n<p>I think that this is some kind of contractive regularization, I need to study more for detailed theoretical background (including the article above:))</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 24962,
      "author_name": "zygmunt",
      "author_url": "",
      "post_date": "05/30/2013 19:29:41",
      "content": "<p></p>\r\n<p>[quote=sayit;24961]</p>\r\n<p>The effect of supervised training unlabeled data with pseudo-labels is that the (neural) network outputs of unlabeled data are closer to 1 or 0 than training only labeled data.</p>\r\n<p>I think that this is some kind of contractive regularization, I need to study more for detailed theoretical background (including the article above:))</p>\r\n<p></p>\r\n<p>[/quote]</p>\r\n<p>Intuitively, a network learns to confirm its own suspicions, so it makes sense.</p>\r\n<p>As far as I understand, the key concept here is a cluster assumption, meaning that each class forms a cluster and the clusters are separated by low density regions. The intro in this paper sums it up nicely:</p>\r\n<p>Chapelle and Zien:&nbsp;Semi-Supervised Classification by Low Density Separation</p>\r\n<p>http://www.kyb.mpg.de/publications/pdfs/pdf2899.pdf</p>\r\n<p></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 25016,
      "author_name": "donghyun",
      "author_url": "",
      "post_date": "06/01/2013 01:51:39",
      "content": "<p><span>In my experiments, </span><strong>pseudo-labels are re-calculated every weights update<strong>&nbsp;in training with labeled and unlabeled data\r\n<strong><strong>simultaneously</strong></strong>.</strong></strong>&nbsp;If we calculate pseudo-label once after training with only labeled data, pseudo-label might be less accurate because the network is overfitted. After training several initial epochs with only\r\n labeled data, the network should be trained with labeled data and unlabeled data using continuously re-calculated pseudo-label. This scheme improve the generalization performance really.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "24931": "",
    "24932": "",
    "24961": "",
    "24962": "",
    "25016": ""
  },
  "source": "meta"
}