{
  "id": 498372,
  "title": "PPO fine-tunes T5 model for generation tasks？",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/498372",
  "author_name": "",
  "post_date": "2024-04-28T03:32:24.868342600Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>As the title shows, I now have the original text and the target text. Since I need to use PPO to fine-tune the T5 model for the generation task, the cross-entropy loss function of the T5 model itself cannot be used because the logits of T5 are no longer returned in the PPO trainer of the TRL library. , its shape is [batch_size, seq_length, hiddeen_size], so cross entropy cannot be calculated, and I hope that the text generated by the T5 model based on the original text should be consistent with the target text, similar to cross entropy training T5, how should I do it? operate?😭</p>",
  "messages": [
    {
      "id": "2780090",
      "postDate": "04/28/2024 03:32:24",
      "content": "<p>As the title shows, I now have the original text and the target text. Since I need to use PPO to fine-tune the T5 model for the generation task, the cross-entropy loss function of the T5 model itself cannot be used because the logits of T5 are no longer returned in the PPO trainer of the TRL library. , its shape is [batch_size, seq_length, hiddeen_size], so cross entropy cannot be calculated, and I hope that the text generated by the T5 model based on the original text should be consistent with the target text, similar to cross entropy training T5, how should I do it? operate?😭</p>",
      "rawMarkdown": "As the title shows, I now have the original text and the target text. Since I need to use PPO to fine-tune the T5 model for the generation task, the cross-entropy loss function of the T5 model itself cannot be used because the logits of T5 are no longer returned in the PPO trainer of the TRL library. , its shape is [batch_size, seq_length, hiddeen_size], so cross entropy cannot be calculated, and I hope that the text generated by the T5 model based on the original text should be consistent with the target text, similar to cross entropy training T5, how should I do it? operate?😭",
      "votes": null
    },
    {
      "id": "2780094",
      "postDate": "04/28/2024 03:38:50",
      "content": "<p>Or, how should the cross entropy loss between the generated text and the target text be regarded as a reward?</p>",
      "rawMarkdown": "Or, how should the cross entropy loss between the generated text and the target text be regarded as a reward?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2780094,
      "author_name": "humbleyll",
      "author_url": "",
      "post_date": "04/28/2024 03:38:50",
      "content": "<p>Or, how should the cross entropy loss between the generated text and the target text be regarded as a reward?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2780090": "As the title shows, I now have the original text and the target text. Since I need to use PPO to fine-tune the T5 model for the generation task, the cross-entropy loss function of the T5 model itself cannot be used because the logits of T5 are no longer returned in the PPO trainer of the TRL library. , its shape is [batch_size, seq_length, hiddeen_size], so cross entropy cannot be calculated, and I hope that the text generated by the T5 model based on the original text should be consistent with the target text, similar to cross entropy training T5, how should I do it? operate?😭",
    "2780094": "Or, how should the cross entropy loss between the generated text and the target text be regarded as a reward?"
  },
  "source": "meta"
}