{
  "id": 457139,
  "title": "How much the 'last dense layer' impacts to the entire result?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/457139",
  "author_name": "",
  "post_date": "2023-11-23T07:13:10.652358400Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'm trying to build a transformer model for this project from scratch for practicing purpose, and I realized while referring other's transformer base line codes that many of them used <strong>a dense layer with 2 dims of output</strong> as the last layer for 2 different experiment types.</p>\n<p>I was trying to build a model which takes inputs containing exp_type information in the sequence (at the first place of the sequence), and expected that the transformer would consider the exp_type information to decide proper inference result with <strong>the last dense layer with only 1 dim of output.</strong></p>\n<p>I'm wondering if kaggle seniors think that the difference of exp_type would be considered enough by a last dense layer weights, and curious about the reason if there's any other answer.</p>\n<p><strong>Are two models' performance similar or different with each other?</strong></p>",
  "messages": [
    {
      "id": "2535220",
      "postDate": "11/23/2023 07:13:10",
      "content": "<p>I'm trying to build a transformer model for this project from scratch for practicing purpose, and I realized while referring other's transformer base line codes that many of them used <strong>a dense layer with 2 dims of output</strong> as the last layer for 2 different experiment types.</p>\n<p>I was trying to build a model which takes inputs containing exp_type information in the sequence (at the first place of the sequence), and expected that the transformer would consider the exp_type information to decide proper inference result with <strong>the last dense layer with only 1 dim of output.</strong></p>\n<p>I'm wondering if kaggle seniors think that the difference of exp_type would be considered enough by a last dense layer weights, and curious about the reason if there's any other answer.</p>\n<p><strong>Are two models' performance similar or different with each other?</strong></p>",
      "rawMarkdown": "I'm trying to build a transformer model for this project from scratch for practicing purpose, and I realized while referring other's transformer base line codes that many of them used **a dense layer with 2 dims of output** as the last layer for 2 different experiment types.\n\nI was trying to build a model which takes inputs containing exp_type information in the sequence (at the first place of the sequence), and expected that the transformer would consider the exp_type information to decide proper inference result with **the last dense layer with only 1 dim of output.**\n\nI'm wondering if kaggle seniors think that the difference of exp_type would be considered enough by a last dense layer weights, and curious about the reason if there's any other answer.\n\n**Are two models' performance similar or different with each other?**",
      "votes": null
    },
    {
      "id": "2535630",
      "postDate": "11/23/2023 13:36:36",
      "content": "<p>I know I have still a lot to learn, but under my intuition one model with two dense layers (one per experiment) will share transformer features and so will be statistically harder to overfit the data. Since they will be relevant for two different predictions.</p>",
      "rawMarkdown": "I know I have still a lot to learn, but under my intuition one model with two dense layers (one per experiment) will share transformer features and so will be statistically harder to overfit the data. Since they will be relevant for two different predictions.",
      "votes": null
    },
    {
      "id": "2535690",
      "postDate": "11/23/2023 14:34:26",
      "content": "<p>Realisitcally you need to experiment. There's many versions. You could also use 1d convs too.</p>",
      "rawMarkdown": "Realisitcally you need to experiment. There's many versions. You could also use 1d convs too.",
      "votes": null
    },
    {
      "id": "2536609",
      "postDate": "11/24/2023 10:02:57",
      "content": "<p>I had tried the two models in small dataset, the performances of the two look similar. Maybe, the performance of the one taking 'exp_type' as input is 'little' better, but it costed almost twice time to train, so I dropped it without taking too much time to experiment it.</p>",
      "rawMarkdown": "I had tried the two models in small dataset, the performances of the two look similar. Maybe, the performance of the one taking 'exp_type' as input is 'little' better, but it costed almost twice time to train, so I dropped it without taking too much time to experiment it.",
      "votes": null
    },
    {
      "id": "2536623",
      "postDate": "11/24/2023 10:19:38",
      "content": "<p>I think as long as you feed the two 'exp_type' into the same transformer, the constraint is still here, so the two models are not much different.</p>",
      "rawMarkdown": "I think as long as you feed the two 'exp_type' into the same transformer, the constraint is still here, so the two models are not much different.",
      "votes": null
    },
    {
      "id": "2536627",
      "postDate": "11/24/2023 10:23:06",
      "content": "<p>Oh, didn't you just put a extra words like '2A3_MaP' and 'DMS_MaP' into your vocab? Why do you think the cost was almost doubled in that method?</p>",
      "rawMarkdown": "Oh, didn't you just put a extra words like '2A3_MaP' and 'DMS_MaP' into your vocab? Why do you think the cost was almost doubled in that method?",
      "votes": null
    },
    {
      "id": "2536664",
      "postDate": "11/24/2023 10:59:26",
      "content": "<p>In such case, the model only outputs one result for '2A3_MaP' or 'DMS_MaP' one time, am I right?</p>",
      "rawMarkdown": "In such case, the model only outputs one result for '2A3_MaP' or 'DMS_MaP' one time, am I right?",
      "votes": null
    },
    {
      "id": "2536684",
      "postDate": "11/24/2023 11:13:32",
      "content": "<p>I think in that case have more flexibility to learn specific features for each 'exp_type'. The loss will be calculated per specific experiment labels. My head hurts @_@</p>\n<p>EDIT: Unless you feed with stratified batches.</p>",
      "rawMarkdown": "I think in that case have more flexibility to learn specific features for each 'exp_type'. The loss will be calculated per specific experiment labels. My head hurts @_@\n\nEDIT: Unless you feed with stratified batches.",
      "votes": null
    },
    {
      "id": "2537252",
      "postDate": "11/25/2023 00:58:42",
      "content": "<p>Yes, because the exp_type might be one of the input words to the transformer model in that case.</p>",
      "rawMarkdown": "Yes, because the exp_type might be one of the input words to the transformer model in that case.",
      "votes": null
    },
    {
      "id": "2537257",
      "postDate": "11/25/2023 01:28:01",
      "content": "<p>As time and resources matter I didn't try yet, but I'm also curious if separated transformer models for each exp_type increase the performance or not. How do you think about that?</p>",
      "rawMarkdown": "As time and resources matter I didn't try yet, but I'm also curious if separated transformer models for each exp_type increase the performance or not. How do you think about that?",
      "votes": null
    },
    {
      "id": "2537298",
      "postDate": "11/25/2023 02:50:31",
      "content": "<p>I think 'separated transformer models for each exp_type' will not increase the performance because it losts the constraint that the two results come from one RNA seq. But I think the two models above we talked about both are 'one transformer for both exp_type', the difference is that, the model with one dim of output will take 'exp_type' as input, and it will output one result depending on the input of 'exp_type' each time; the model with two dims of output will output two results for two exp_type simultaneously. I think these two models are not much different, as I know there is no absolute answer for this. As I mentioned above, for this competition, maybe, the first one is slightly better, but the conclusion is not based on enough experiment.</p>",
      "rawMarkdown": "I think 'separated transformer models for each exp_type' will not increase the performance because it losts the constraint that the two results come from one RNA seq. But I think the two models above we talked about both are 'one transformer for both exp_type', the difference is that, the model with one dim of output will take 'exp_type' as input, and it will output one result depending on the input of 'exp_type' each time; the model with two dims of output will output two results for two exp_type simultaneously. I think these two models are not much different, as I know there is no absolute answer for this. As I mentioned above, for this competition, maybe, the first one is slightly better, but the conclusion is not based on enough experiment.",
      "votes": null
    },
    {
      "id": "2537313",
      "postDate": "11/25/2023 03:39:53",
      "content": "<p>Sure, I get it. Thank you for your reply :)</p>",
      "rawMarkdown": "Sure, I get it. Thank you for your reply :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2535630,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/23/2023 13:36:36",
      "content": "<p>I know I have still a lot to learn, but under my intuition one model with two dense layers (one per experiment) will share transformer features and so will be statistically harder to overfit the data. Since they will be relevant for two different predictions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2536623,
          "author_name": "hermitmb",
          "author_url": "",
          "post_date": "11/24/2023 10:19:38",
          "content": "<p>I think as long as you feed the two 'exp_type' into the same transformer, the constraint is still here, so the two models are not much different.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2536684,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "11/24/2023 11:13:32",
              "content": "<p>I think in that case have more flexibility to learn specific features for each 'exp_type'. The loss will be calculated per specific experiment labels. My head hurts @_@</p>\n<p>EDIT: Unless you feed with stratified batches.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2535690,
      "author_name": "themecheng",
      "author_url": "",
      "post_date": "11/23/2023 14:34:26",
      "content": "<p>Realisitcally you need to experiment. There's many versions. You could also use 1d convs too.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2536609,
      "author_name": "hermitmb",
      "author_url": "",
      "post_date": "11/24/2023 10:02:57",
      "content": "<p>I had tried the two models in small dataset, the performances of the two look similar. Maybe, the performance of the one taking 'exp_type' as input is 'little' better, but it costed almost twice time to train, so I dropped it without taking too much time to experiment it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2536627,
          "author_name": "taegeunlim",
          "author_url": "",
          "post_date": "11/24/2023 10:23:06",
          "content": "<p>Oh, didn't you just put a extra words like '2A3_MaP' and 'DMS_MaP' into your vocab? Why do you think the cost was almost doubled in that method?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2536664,
              "author_name": "hermitmb",
              "author_url": "",
              "post_date": "11/24/2023 10:59:26",
              "content": "<p>In such case, the model only outputs one result for '2A3_MaP' or 'DMS_MaP' one time, am I right?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2537252,
                  "author_name": "taegeunlim",
                  "author_url": "",
                  "post_date": "11/25/2023 00:58:42",
                  "content": "<p>Yes, because the exp_type might be one of the input words to the transformer model in that case.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2537257,
          "author_name": "taegeunlim",
          "author_url": "",
          "post_date": "11/25/2023 01:28:01",
          "content": "<p>As time and resources matter I didn't try yet, but I'm also curious if separated transformer models for each exp_type increase the performance or not. How do you think about that?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2537298,
              "author_name": "hermitmb",
              "author_url": "",
              "post_date": "11/25/2023 02:50:31",
              "content": "<p>I think 'separated transformer models for each exp_type' will not increase the performance because it losts the constraint that the two results come from one RNA seq. But I think the two models above we talked about both are 'one transformer for both exp_type', the difference is that, the model with one dim of output will take 'exp_type' as input, and it will output one result depending on the input of 'exp_type' each time; the model with two dims of output will output two results for two exp_type simultaneously. I think these two models are not much different, as I know there is no absolute answer for this. As I mentioned above, for this competition, maybe, the first one is slightly better, but the conclusion is not based on enough experiment.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2537313,
                  "author_name": "taegeunlim",
                  "author_url": "",
                  "post_date": "11/25/2023 03:39:53",
                  "content": "<p>Sure, I get it. Thank you for your reply :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2535220": "I'm trying to build a transformer model for this project from scratch for practicing purpose, and I realized while referring other's transformer base line codes that many of them used **a dense layer with 2 dims of output** as the last layer for 2 different experiment types.\n\nI was trying to build a model which takes inputs containing exp_type information in the sequence (at the first place of the sequence), and expected that the transformer would consider the exp_type information to decide proper inference result with **the last dense layer with only 1 dim of output.**\n\nI'm wondering if kaggle seniors think that the difference of exp_type would be considered enough by a last dense layer weights, and curious about the reason if there's any other answer.\n\n**Are two models' performance similar or different with each other?**",
    "2535630": "I know I have still a lot to learn, but under my intuition one model with two dense layers (one per experiment) will share transformer features and so will be statistically harder to overfit the data. Since they will be relevant for two different predictions.",
    "2535690": "Realisitcally you need to experiment. There's many versions. You could also use 1d convs too.",
    "2536609": "I had tried the two models in small dataset, the performances of the two look similar. Maybe, the performance of the one taking 'exp_type' as input is 'little' better, but it costed almost twice time to train, so I dropped it without taking too much time to experiment it.",
    "2536623": "I think as long as you feed the two 'exp_type' into the same transformer, the constraint is still here, so the two models are not much different.",
    "2536627": "Oh, didn't you just put a extra words like '2A3_MaP' and 'DMS_MaP' into your vocab? Why do you think the cost was almost doubled in that method?",
    "2536664": "In such case, the model only outputs one result for '2A3_MaP' or 'DMS_MaP' one time, am I right?",
    "2536684": "I think in that case have more flexibility to learn specific features for each 'exp_type'. The loss will be calculated per specific experiment labels. My head hurts @_@\n\nEDIT: Unless you feed with stratified batches.",
    "2537252": "Yes, because the exp_type might be one of the input words to the transformer model in that case.",
    "2537257": "As time and resources matter I didn't try yet, but I'm also curious if separated transformer models for each exp_type increase the performance or not. How do you think about that?",
    "2537298": "I think 'separated transformer models for each exp_type' will not increase the performance because it losts the constraint that the two results come from one RNA seq. But I think the two models above we talked about both are 'one transformer for both exp_type', the difference is that, the model with one dim of output will take 'exp_type' as input, and it will output one result depending on the input of 'exp_type' each time; the model with two dims of output will output two results for two exp_type simultaneously. I think these two models are not much different, as I know there is no absolute answer for this. As I mentioned above, for this competition, maybe, the first one is slightly better, but the conclusion is not based on enough experiment.",
    "2537313": "Sure, I get it. Thank you for your reply :)"
  },
  "source": "meta"
}