{
  "id": 78044,
  "title": "Reinforcement Learning for NLP",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78044",
  "author_name": "",
  "post_date": "2019-01-19T05:08:38.977895600Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Has anyone tried to use Reinforcement Learning for NLP? Can RL be  used in this case of predicting toxic content? I was thinking the output of LSTM  could be used as input to a Deep Q-Learning network. What could be the reward design? Any thoughts on that? </p>",
  "messages": [
    {
      "id": "458211",
      "postDate": "01/19/2019 05:08:38",
      "content": "<p>Has anyone tried to use Reinforcement Learning for NLP? Can RL be  used in this case of predicting toxic content? I was thinking the output of LSTM  could be used as input to a Deep Q-Learning network. What could be the reward design? Any thoughts on that? </p>",
      "rawMarkdown": "Has anyone tried to use Reinforcement Learning for NLP? Can RL be  used in this case of predicting toxic content? I was thinking the output of LSTM  could be used as input to a Deep Q-Learning network. What could be the reward design? Any thoughts on that?",
      "votes": null
    },
    {
      "id": "459955",
      "postDate": "01/22/2019 17:37:10",
      "content": "<p>I created a kernel to explore RL - deep Q-Learning for predicting toxic/non-toxic comment. Here is the kernel - <a href=\"https://www.kaggle.com/amitabhac/rl-nlp-first-attempt\">https://www.kaggle.com/amitabhac/rl-nlp-first-attempt</a>. Even though training is encouraging the prediction is not good. It is predicting all zeroes. Please take a look and let me know what you think.</p>",
      "rawMarkdown": "I created a kernel to explore RL - deep Q-Learning for predicting toxic/non-toxic comment. Here is the kernel - https://www.kaggle.com/amitabhac/rl-nlp-first-attempt. Even though training is encouraging the prediction is not good. It is predicting all zeroes. Please take a look and let me know what you think.",
      "votes": null
    },
    {
      "id": "460088",
      "postDate": "01/23/2019 01:25:38",
      "content": "<p>RL Learning is notorious for how computationally expensive it is. So it seems to me that it would be kind of unrealistic to get good performance out of it with the time constraints.</p>",
      "rawMarkdown": "RL Learning is notorious for how computationally expensive it is. So it seems to me that it would be kind of unrealistic to get good performance out of it with the time constraints.",
      "votes": null
    },
    {
      "id": "460105",
      "postDate": "01/23/2019 02:15:44",
      "content": "<p>It is unrealistic to use RL in this competition given the required number of hours to train a model as per rule and agree on @MitchelFung point of view. Do not fall on trap that any fancy algorithm/s can solve everything. The key success on this competition are: comprehensive Pre-processing, highly-optimize model architecture and some post-processing (aka secret sauce)..</p>",
      "rawMarkdown": "It is unrealistic to use RL in this competition given the required number of hours to train a model as per rule and agree on @MitchelFung point of view. Do not fall on trap that any fancy algorithm/s can solve everything. The key success on this competition are: comprehensive Pre-processing, highly-optimize model architecture and some post-processing (aka secret sauce)..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 459955,
      "author_name": "amitabhac",
      "author_url": "",
      "post_date": "01/22/2019 17:37:10",
      "content": "<p>I created a kernel to explore RL - deep Q-Learning for predicting toxic/non-toxic comment. Here is the kernel - <a href=\"https://www.kaggle.com/amitabhac/rl-nlp-first-attempt\">https://www.kaggle.com/amitabhac/rl-nlp-first-attempt</a>. Even though training is encouraging the prediction is not good. It is predicting all zeroes. Please take a look and let me know what you think.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460088,
      "author_name": "msf908",
      "author_url": "",
      "post_date": "01/23/2019 01:25:38",
      "content": "<p>RL Learning is notorious for how computationally expensive it is. So it seems to me that it would be kind of unrealistic to get good performance out of it with the time constraints.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460105,
      "author_name": "",
      "author_url": "",
      "post_date": "01/23/2019 02:15:44",
      "content": "<p>It is unrealistic to use RL in this competition given the required number of hours to train a model as per rule and agree on @MitchelFung point of view. Do not fall on trap that any fancy algorithm/s can solve everything. The key success on this competition are: comprehensive Pre-processing, highly-optimize model architecture and some post-processing (aka secret sauce)..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "458211": "Has anyone tried to use Reinforcement Learning for NLP? Can RL be  used in this case of predicting toxic content? I was thinking the output of LSTM  could be used as input to a Deep Q-Learning network. What could be the reward design? Any thoughts on that?",
    "459955": "I created a kernel to explore RL - deep Q-Learning for predicting toxic/non-toxic comment. Here is the kernel - https://www.kaggle.com/amitabhac/rl-nlp-first-attempt. Even though training is encouraging the prediction is not good. It is predicting all zeroes. Please take a look and let me know what you think.",
    "460088": "RL Learning is notorious for how computationally expensive it is. So it seems to me that it would be kind of unrealistic to get good performance out of it with the time constraints.",
    "460105": "It is unrealistic to use RL in this competition given the required number of hours to train a model as per rule and agree on @MitchelFung point of view. Do not fall on trap that any fancy algorithm/s can solve everything. The key success on this competition are: comprehensive Pre-processing, highly-optimize model architecture and some post-processing (aka secret sauce).."
  },
  "source": "meta"
}