{
  "id": 93734,
  "title": "Reverse Learning",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/93734",
  "author_name": "Granddad",
  "post_date": "2019-05-29T14:35:24.516000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Here's an idea that we probably won't have time to implement before the contest closes.  We want to post it anyway in case someone else finds this useful or interesting.</p>\n\n<p>In what follows, \"training data\" will refer to the noisy, inaccurately labeled data, \"validation data\" will refer to the curated, accurately labeled data, and \"testing data\" will refer to the unlabeled data.</p>\n\n<p>So the idea is to adjust the inaccurate training labels in order to make them better predictors.  We can judge the predictive power of our training labels by measuring success on the validation data.  There may be several ways to implement this.  Here is one.</p>\n\n<p>Train a neural net the usual way using training data.  As it is training, adjust the weights and biases (model parameters) in the preferred direction (the direction of the negative gradient of the training loss function).  The validation data is typically never back-tracked.  But if it was, it would also show a preferred direction to push model parameters.  The idea is to use this validation direction to make adjustments in the training labels.</p>\n\n<p>To be more precise, suppose for the sake of simplicity there is only one label and we need to decide \"Yes\" or \"No\" for each training input.  The validation data would like to push the model parameters in a preferred direction, V.   Similarly, each training input, i,  has two possible preferred directions, Y_i and N_i, depending on its label.  It can be shown that Y_i and N_i are directly opposite (assuming a symmetry condition on the loss function).  So exactly one of them will be within 90 degrees of V and hence put negative pressure on the validation loss.  In this way, the validation data \"suggests\" a label for this input.  This suggested label may change as we keep training.  So at each batch, we move the input labels toward what is suggested by the validation data.</p>\n\n<p>Assuming this makes sense, there are issues that need to be resolved through experimentation.  How do we minimize the work of backtracking with 80 different labels?  Can we get good label suggestions by backtracking through just a few layers? Do we move all labels slightly toward their suggested label? Or flip a few of the labels completely?  How often should we restart the training process with the updated labels? Should there be an ensemble of learners simultaneously voting on labels?  Etc.</p>\n\n<p>\"Training data\" and \"validation data\" might be backward from what is expected, as most researchers would prefer training on the most accurate data to adjust labels on the inaccurate\ndata.  But the reverse approach can be philosophically justified two ways.  First, consider that\nwhen a human learns, he/she will directly apply prior knowledge to reach new conclusions.   The reverse is just as important.  That is, consider what implications new conclusions would have on\nprior knowledge and adjust these conclusions accordingly.   The human, trying to make sense of the world, is continually looking for consistency.</p>\n\n<p>The second justification is more pragmatic.  How is new knowledge going to be used?  The reason for improving training label accuracy is not because we are interested in the labels per se.  Rather it is so we can use these labels to accurately predict unlabeled test data.  So why not adjust the labels to directly fulfill this role?  If an accordion sound clip sounds like a harmonica, it might possibly give better predictions if it were re-labeled a harmonica, or perhaps half-accordion and half-harmonica.</p>\n\n<p>Please feel free to leave comments, references, etc.</p>\n\n<p>--Eric and Granddad\nSound Learning Team</p>",
  "messages": [
    {
      "id": 539110,
      "postDate": "2019-05-29T14:35:24.517Z",
      "content": "<p>Here's an idea that we probably won't have time to implement before the contest closes.  We want to post it anyway in case someone else finds this useful or interesting.</p>\n\n<p>In what follows, \"training data\" will refer to the noisy, inaccurately labeled data, \"validation data\" will refer to the curated, accurately labeled data, and \"testing data\" will refer to the unlabeled data.</p>\n\n<p>So the idea is to adjust the inaccurate training labels in order to make them better predictors.  We can judge the predictive power of our training labels by measuring success on the validation data.  There may be several ways to implement this.  Here is one.</p>\n\n<p>Train a neural net the usual way using training data.  As it is training, adjust the weights and biases (model parameters) in the preferred direction (the direction of the negative gradient of the training loss function).  The validation data is typically never back-tracked.  But if it was, it would also show a preferred direction to push model parameters.  The idea is to use this validation direction to make adjustments in the training labels.</p>\n\n<p>To be more precise, suppose for the sake of simplicity there is only one label and we need to decide \"Yes\" or \"No\" for each training input.  The validation data would like to push the model parameters in a preferred direction, V.   Similarly, each training input, i,  has two possible preferred directions, Y_i and N_i, depending on its label.  It can be shown that Y_i and N_i are directly opposite (assuming a symmetry condition on the loss function).  So exactly one of them will be within 90 degrees of V and hence put negative pressure on the validation loss.  In this way, the validation data \"suggests\" a label for this input.  This suggested label may change as we keep training.  So at each batch, we move the input labels toward what is suggested by the validation data.</p>\n\n<p>Assuming this makes sense, there are issues that need to be resolved through experimentation.  How do we minimize the work of backtracking with 80 different labels?  Can we get good label suggestions by backtracking through just a few layers? Do we move all labels slightly toward their suggested label? Or flip a few of the labels completely?  How often should we restart the training process with the updated labels? Should there be an ensemble of learners simultaneously voting on labels?  Etc.</p>\n\n<p>\"Training data\" and \"validation data\" might be backward from what is expected, as most researchers would prefer training on the most accurate data to adjust labels on the inaccurate\ndata.  But the reverse approach can be philosophically justified two ways.  First, consider that\nwhen a human learns, he/she will directly apply prior knowledge to reach new conclusions.   The reverse is just as important.  That is, consider what implications new conclusions would have on\nprior knowledge and adjust these conclusions accordingly.   The human, trying to make sense of the world, is continually looking for consistency.</p>\n\n<p>The second justification is more pragmatic.  How is new knowledge going to be used?  The reason for improving training label accuracy is not because we are interested in the labels per se.  Rather it is so we can use these labels to accurately predict unlabeled test data.  So why not adjust the labels to directly fulfill this role?  If an accordion sound clip sounds like a harmonica, it might possibly give better predictions if it were re-labeled a harmonica, or perhaps half-accordion and half-harmonica.</p>\n\n<p>Please feel free to leave comments, references, etc.</p>\n\n<p>--Eric and Granddad\nSound Learning Team</p>",
      "rawMarkdown": "Here's an idea that we probably won't have time to implement before the contest closes.  We want to post it anyway in case someone else finds this useful or interesting.\n\nIn what follows, \"training data\" will refer to the noisy, inaccurately labeled data, \"validation data\" will refer to the curated, accurately labeled data, and \"testing data\" will refer to the unlabeled data.\n\nSo the idea is to adjust the inaccurate training labels in order to make them better predictors.  We can judge the predictive power of our training labels by measuring success on the validation data.  There may be several ways to implement this.  Here is one.\n\nTrain a neural net the usual way using training data.  As it is training, adjust the weights and biases (model parameters) in the preferred direction (the direction of the negative gradient of the training loss function).  The validation data is typically never back-tracked.  But if it was, it would also show a preferred direction to push model parameters.  The idea is to use this validation direction to make adjustments in the training labels.\n\nTo be more precise, suppose for the sake of simplicity there is only one label and we need to decide \"Yes\" or \"No\" for each training input.  The validation data would like to push the model parameters in a preferred direction, V.   Similarly, each training input, i,  has two possible preferred directions, Y_i and N_i, depending on its label.  It can be shown that Y_i and N_i are directly opposite (assuming a symmetry condition on the loss function).  So exactly one of them will be within 90 degrees of V and hence put negative pressure on the validation loss.  In this way, the validation data \"suggests\" a label for this input.  This suggested label may change as we keep training.  So at each batch, we move the input labels toward what is suggested by the validation data.\n\nAssuming this makes sense, there are issues that need to be resolved through experimentation.  How do we minimize the work of backtracking with 80 different labels?  Can we get good label suggestions by backtracking through just a few layers? Do we move all labels slightly toward their suggested label? Or flip a few of the labels completely?  How often should we restart the training process with the updated labels? Should there be an ensemble of learners simultaneously voting on labels?  Etc.\n\n\"Training data\" and \"validation data\" might be backward from what is expected, as most researchers would prefer training on the most accurate data to adjust labels on the inaccurate\ndata.  But the reverse approach can be philosophically justified two ways.  First, consider that\nwhen a human learns, he/she will directly apply prior knowledge to reach new conclusions.   The reverse is just as important.  That is, consider what implications new conclusions would have on\nprior knowledge and adjust these conclusions accordingly.   The human, trying to make sense of the world, is continually looking for consistency.\n\nThe second justification is more pragmatic.  How is new knowledge going to be used?  The reason for improving training label accuracy is not because we are interested in the labels per se.  Rather it is so we can use these labels to accurately predict unlabeled test data.  So why not adjust the labels to directly fulfill this role?  If an accordion sound clip sounds like a harmonica, it might possibly give better predictions if it were re-labeled a harmonica, or perhaps half-accordion and half-harmonica.\n\nPlease feel free to leave comments, references, etc.\n\n--Eric and Granddad\nSound Learning Team\n",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "539110": "Here's an idea that we probably won't have time to implement before the contest closes.  We want to post it anyway in case someone else finds this useful or interesting.\n\nIn what follows, \"training data\" will refer to the noisy, inaccurately labeled data, \"validation data\" will refer to the curated, accurately labeled data, and \"testing data\" will refer to the unlabeled data.\n\nSo the idea is to adjust the inaccurate training labels in order to make them better predictors.  We can judge the predictive power of our training labels by measuring success on the validation data.  There may be several ways to implement this.  Here is one.\n\nTrain a neural net the usual way using training data.  As it is training, adjust the weights and biases (model parameters) in the preferred direction (the direction of the negative gradient of the training loss function).  The validation data is typically never back-tracked.  But if it was, it would also show a preferred direction to push model parameters.  The idea is to use this validation direction to make adjustments in the training labels.\n\nTo be more precise, suppose for the sake of simplicity there is only one label and we need to decide \"Yes\" or \"No\" for each training input.  The validation data would like to push the model parameters in a preferred direction, V.   Similarly, each training input, i,  has two possible preferred directions, Y_i and N_i, depending on its label.  It can be shown that Y_i and N_i are directly opposite (assuming a symmetry condition on the loss function).  So exactly one of them will be within 90 degrees of V and hence put negative pressure on the validation loss.  In this way, the validation data \"suggests\" a label for this input.  This suggested label may change as we keep training.  So at each batch, we move the input labels toward what is suggested by the validation data.\n\nAssuming this makes sense, there are issues that need to be resolved through experimentation.  How do we minimize the work of backtracking with 80 different labels?  Can we get good label suggestions by backtracking through just a few layers? Do we move all labels slightly toward their suggested label? Or flip a few of the labels completely?  How often should we restart the training process with the updated labels? Should there be an ensemble of learners simultaneously voting on labels?  Etc.\n\n\"Training data\" and \"validation data\" might be backward from what is expected, as most researchers would prefer training on the most accurate data to adjust labels on the inaccurate\ndata.  But the reverse approach can be philosophically justified two ways.  First, consider that\nwhen a human learns, he/she will directly apply prior knowledge to reach new conclusions.   The reverse is just as important.  That is, consider what implications new conclusions would have on\nprior knowledge and adjust these conclusions accordingly.   The human, trying to make sense of the world, is continually looking for consistency.\n\nThe second justification is more pragmatic.  How is new knowledge going to be used?  The reason for improving training label accuracy is not because we are interested in the labels per se.  Rather it is so we can use these labels to accurately predict unlabeled test data.  So why not adjust the labels to directly fulfill this role?  If an accordion sound clip sounds like a harmonica, it might possibly give better predictions if it were re-labeled a harmonica, or perhaps half-accordion and half-harmonica.\n\nPlease feel free to leave comments, references, etc.\n\n--Eric and Granddad\nSound Learning Team\n"
  }
}