{
  "id": 433906,
  "title": "Adding validation data (only 1000 samples) into training makes training process much slower",
  "url": "/competitions/asl-fingerspelling/discussion/433906",
  "author_name": "",
  "post_date": "2023-08-23T10:36:09.982248200Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi guys,</p>\n<p>Previously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.</p>\n<p>However, I found adding these 1000 samples into training significantly slows the convergence speed.</p>\n<p>With the validation set, the training score ends up with ~0.86.</p>\n<p>However, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).</p>\n<p>It seems to me I need to increase 1/5~1/6 epochs for training using all data. </p>\n<p>I'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 </p>\n<p>Is this normal? </p>",
  "messages": [
    {
      "id": "2404584",
      "postDate": "08/23/2023 10:36:09",
      "content": "<p>Hi guys,</p>\n<p>Previously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.</p>\n<p>However, I found adding these 1000 samples into training significantly slows the convergence speed.</p>\n<p>With the validation set, the training score ends up with ~0.86.</p>\n<p>However, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).</p>\n<p>It seems to me I need to increase 1/5~1/6 epochs for training using all data. </p>\n<p>I'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 </p>\n<p>Is this normal? </p>",
      "rawMarkdown": "Hi guys,\n\nPreviously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.\n\nHowever, I found adding these 1000 samples into training significantly slows the convergence speed.\n\nWith the validation set, the training score ends up with ~0.86.\n\nHowever, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).\n\n It seems to me I need to increase 1/5~1/6 epochs for training using all data. \n\nI'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 \n\nIs this normal?",
      "votes": null
    },
    {
      "id": "2404651",
      "postDate": "08/23/2023 12:17:53",
      "content": "<p>use previous probability as regularization:</p>\n<pre><code>loss( sample) = ctc_loss_or other + \n</code></pre>",
      "rawMarkdown": "use previous probability as regularization:\n```\nloss(new sample) = ctc_loss_or other_loss(new sample|truth phrase) + KLDiv(old_probability, current_probability)\n\n```",
      "votes": null
    },
    {
      "id": "2404677",
      "postDate": "08/23/2023 12:42:08",
      "content": "<p>Thank you for your reply!</p>\n<p>what's the definition of probability here?</p>",
      "rawMarkdown": "Thank you for your reply!\n\nwhat's the definition of probability here?",
      "votes": null
    },
    {
      "id": "2405547",
      "postDate": "08/24/2023 01:23:44",
      "content": "<p>I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch</p>",
      "rawMarkdown": "I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch",
      "votes": null
    },
    {
      "id": "2405669",
      "postDate": "08/24/2023 03:53:39",
      "content": "<p>Thanks for your reply! It seems make sense</p>",
      "rawMarkdown": "Thanks for your reply! It seems make sense",
      "votes": null
    },
    {
      "id": "2405683",
      "postDate": "08/24/2023 04:09:51",
      "content": "<p>the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.</p>\n<p>old_prob is from previous model.</p>",
      "rawMarkdown": "the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.\n\nold_prob is from previous model.",
      "votes": null
    },
    {
      "id": "2405684",
      "postDate": "08/24/2023 04:10:52",
      "content": "<p>or you can use uniform probability for old_prob</p>\n<p><a href=\"https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\" target=\"_blank\">https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392</a><br>\n<a href=\"https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py\" target=\"_blank\">https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py</a></p>",
      "rawMarkdown": "or you can use uniform probability for old_prob\n\nhttps://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\nhttps://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py",
      "votes": null
    },
    {
      "id": "2405691",
      "postDate": "08/24/2023 04:17:14",
      "content": "<p>\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"<br>\nif prob_old is from a model, the model is acting as a teaching.<br>\nif prob_old is from a assumption(e.g. uniform prob), it is a prior </p>\n<p><a href=\"https://arxiv.org/pdf/2005.09310.pdf\" target=\"_blank\">https://arxiv.org/pdf/2005.09310.pdf</a></p>",
      "rawMarkdown": "\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"\nif prob_old is from a model, the model is acting as a teaching.\nif prob_old is from a assumption(e.g. uniform prob), it is a prior \n\nhttps://arxiv.org/pdf/2005.09310.pdf",
      "votes": null
    },
    {
      "id": "2406019",
      "postDate": "08/24/2023 07:34:43",
      "content": "<p>looks so advanced. don't know if I have time to try it out at the last minute…</p>",
      "rawMarkdown": "looks so advanced. don't know if I have time to try it out at the last minute...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2404651,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/23/2023 12:17:53",
      "content": "<p>use previous probability as regularization:</p>\n<pre><code>loss( sample) = ctc_loss_or other + \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2404677,
          "author_name": "nightsh4de",
          "author_url": "",
          "post_date": "08/23/2023 12:42:08",
          "content": "<p>Thank you for your reply!</p>\n<p>what's the definition of probability here?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2405547,
              "author_name": "coldfir3",
              "author_url": "",
              "post_date": "08/24/2023 01:23:44",
              "content": "<p>I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2405669,
                  "author_name": "nightsh4de",
                  "author_url": "",
                  "post_date": "08/24/2023 03:53:39",
                  "content": "<p>Thanks for your reply! It seems make sense</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2405683,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "08/24/2023 04:09:51",
                      "content": "<p>the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.</p>\n<p>old_prob is from previous model.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2405684,
                          "author_name": "hengck23",
                          "author_url": "",
                          "post_date": "08/24/2023 04:10:52",
                          "content": "<p>or you can use uniform probability for old_prob</p>\n<p><a href=\"https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\" target=\"_blank\">https://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392</a><br>\n<a href=\"https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py\" target=\"_blank\">https://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py</a></p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2405691,
                              "author_name": "hengck23",
                              "author_url": "",
                              "post_date": "08/24/2023 04:17:14",
                              "content": "<p>\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"<br>\nif prob_old is from a model, the model is acting as a teaching.<br>\nif prob_old is from a assumption(e.g. uniform prob), it is a prior </p>\n<p><a href=\"https://arxiv.org/pdf/2005.09310.pdf\" target=\"_blank\">https://arxiv.org/pdf/2005.09310.pdf</a></p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2406019,
                                  "author_name": "nightsh4de",
                                  "author_url": "",
                                  "post_date": "08/24/2023 07:34:43",
                                  "content": "<p>looks so advanced. don't know if I have time to try it out at the last minute…</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2404584": "Hi guys,\n\nPreviously, I used 234418913.parquet (1000 samples) as the validation set, and used the remaining data as the training set.\n\nHowever, I found adding these 1000 samples into training significantly slows the convergence speed.\n\nWith the validation set, the training score ends up with ~0.86.\n\nHowever, using full data ends up with ~0.84 training score under exactly the same setup (LB score also decreased).\n\n It seems to me I need to increase 1/5~1/6 epochs for training using all data. \n\nI'm so confused why 1000 samples (only 1/67 of all data) can make such a big difference. 🤯 \n\nIs this normal?",
    "2404651": "use previous probability as regularization:\n```\nloss(new sample) = ctc_loss_or other_loss(new sample|truth phrase) + KLDiv(old_probability, current_probability)\n\n```",
    "2404677": "Thank you for your reply!\n\nwhat's the definition of probability here?",
    "2405547": "I think he means something like the softmax of the logits there. I belive old_prob is from the past epoch and current_prob is for the current epoch",
    "2405669": "Thanks for your reply! It seems make sense",
    "2405683": "the objective is to make the prediction close to the previous model (i.e. retain previous model accuracy) when possible.\n\nold_prob is from previous model.",
    "2405684": "or you can use uniform probability for old_prob\n\nhttps://discuss.pytorch.org/t/label-smoothing-with-ctcloss/103392\nhttps://github.com/TeaPoly/CTC-OptimizedLoss/blob/main/ctc_label_smoothing_loss.py",
    "2405691": "\"From Fig. 2, it becomes clear that KD acts as an informed label smoothing,\"\nif prob_old is from a model, the model is acting as a teaching.\nif prob_old is from a assumption(e.g. uniform prob), it is a prior \n\nhttps://arxiv.org/pdf/2005.09310.pdf",
    "2406019": "looks so advanced. don't know if I have time to try it out at the last minute..."
  },
  "source": "meta"
}