{
  "id": 313864,
  "title": "How to pass boss line?",
  "url": "/competitions/ml2022spring-hw4/discussion/313864",
  "author_name": "",
  "post_date": "2022-03-19T10:07:28.096284300Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have used all the technics mentioned in \"https://speech.ee.ntu.edu.tw/~hylee/ml/ml2022-course-data/Machine%20Learning%20HW4.pdf\", but I still can't reach the boss line.</p>\n<p>I think the main reason is in the hyperparameters. I reckon that the width of feed forward network inside conformer is very important, so I set it  ffn_dim=1024, while num_layers=6(because I think more layers are harder to train)</p>\n<p>The depthwise_conv_kernel_size is set to 31 refer to the original paper. </p>\n<p>So, how to achieve boss line? Would you like to share your tricks with me?</p>",
  "messages": [
    {
      "id": "1728818",
      "postDate": "03/19/2022 10:07:28",
      "content": "<p>I have used all the technics mentioned in \"https://speech.ee.ntu.edu.tw/~hylee/ml/ml2022-course-data/Machine%20Learning%20HW4.pdf\", but I still can't reach the boss line.</p>\n<p>I think the main reason is in the hyperparameters. I reckon that the width of feed forward network inside conformer is very important, so I set it  ffn_dim=1024, while num_layers=6(because I think more layers are harder to train)</p>\n<p>The depthwise_conv_kernel_size is set to 31 refer to the original paper. </p>\n<p>So, how to achieve boss line? Would you like to share your tricks with me?</p>",
      "rawMarkdown": "I have used all the technics mentioned in \"https://speech.ee.ntu.edu.tw/~hylee/ml/ml2022-course-data/Machine%20Learning%20HW4.pdf\", but I still can't reach the boss line.\n\nI think the main reason is in the hyperparameters. I reckon that the width of feed forward network inside conformer is very important, so I set it  ffn_dim=1024, while num_layers=6(because I think more layers are harder to train)\n\nThe depthwise_conv_kernel_size is set to 31 refer to the original paper. \n\nSo, how to achieve boss line? Would you like to share your tricks with me?",
      "votes": null
    },
    {
      "id": "1729591",
      "postDate": "03/20/2022 09:06:22",
      "content": "<p>Besides…. My validation accuracy reached 0.96, but test is much lower.</p>\n<p><a href=\"https://ibb.co/z6Zt2J9\" target=\"_blank\">https://ibb.co/z6Zt2J9</a></p>",
      "rawMarkdown": "Besides.... My validation accuracy reached 0.96, but test is much lower.\n\nhttps://ibb.co/z6Zt2J9",
      "votes": null
    },
    {
      "id": "1729598",
      "postDate": "03/20/2022 09:14:19",
      "content": "<p>HI, <br>\nDid you implement attention pooling &amp; AMSoftmax?<br>\nSince using AMSoftmax may cause training to be harder, my suggestion is that trying to adjust the <code>m</code> value in the formula smaller (the original value suggested in the paper is about 0.3x, but in this homework it can be too large and thus causing the training of the model becomes header.).</p>",
      "rawMarkdown": "HI, \nDid you implement attention pooling & AMSoftmax?\nSince using AMSoftmax may cause training to be harder, my suggestion is that trying to adjust the `m` value in the formula smaller (the original value suggested in the paper is about 0.3x, but in this homework it can be too large and thus causing the training of the model becomes header.).",
      "votes": null
    },
    {
      "id": "1729600",
      "postDate": "03/20/2022 09:17:35",
      "content": "<p>It's overfitting problem. In this homework, the accuracy on validation set and test set may have a gap. But I think gap between 0.96 and 0.82 is too large.</p>",
      "rawMarkdown": "It's overfitting problem. In this homework, the accuracy on validation set and test set may have a gap. But I think gap between 0.96 and 0.82 is too large.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1729591,
      "author_name": "andrewchenchen",
      "author_url": "",
      "post_date": "03/20/2022 09:06:22",
      "content": "<p>Besides…. My validation accuracy reached 0.96, but test is much lower.</p>\n<p><a href=\"https://ibb.co/z6Zt2J9\" target=\"_blank\">https://ibb.co/z6Zt2J9</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1729600,
          "author_name": "b08902126hanklin",
          "author_url": "",
          "post_date": "03/20/2022 09:17:35",
          "content": "<p>It's overfitting problem. In this homework, the accuracy on validation set and test set may have a gap. But I think gap between 0.96 and 0.82 is too large.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1729598,
      "author_name": "b08902126hanklin",
      "author_url": "",
      "post_date": "03/20/2022 09:14:19",
      "content": "<p>HI, <br>\nDid you implement attention pooling &amp; AMSoftmax?<br>\nSince using AMSoftmax may cause training to be harder, my suggestion is that trying to adjust the <code>m</code> value in the formula smaller (the original value suggested in the paper is about 0.3x, but in this homework it can be too large and thus causing the training of the model becomes header.).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1728818": "I have used all the technics mentioned in \"https://speech.ee.ntu.edu.tw/~hylee/ml/ml2022-course-data/Machine%20Learning%20HW4.pdf\", but I still can't reach the boss line.\n\nI think the main reason is in the hyperparameters. I reckon that the width of feed forward network inside conformer is very important, so I set it  ffn_dim=1024, while num_layers=6(because I think more layers are harder to train)\n\nThe depthwise_conv_kernel_size is set to 31 refer to the original paper. \n\nSo, how to achieve boss line? Would you like to share your tricks with me?",
    "1729591": "Besides.... My validation accuracy reached 0.96, but test is much lower.\n\nhttps://ibb.co/z6Zt2J9",
    "1729598": "HI, \nDid you implement attention pooling & AMSoftmax?\nSince using AMSoftmax may cause training to be harder, my suggestion is that trying to adjust the `m` value in the formula smaller (the original value suggested in the paper is about 0.3x, but in this homework it can be too large and thus causing the training of the model becomes header.).",
    "1729600": "It's overfitting problem. In this homework, the accuracy on validation set and test set may have a gap. But I think gap between 0.96 and 0.82 is too large."
  },
  "source": "meta"
}