{
  "id": 241557,
  "title": "Need some guide in building a simple transformer Model baseline",
  "url": "/competitions/bms-molecular-translation/discussion/241557",
  "author_name": "",
  "post_date": "2021-05-25T04:47:00.225797800Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is my simple efficientnet-b2 + transformer model <br>\nit seems hardly handle with training loss<br>\n<a href=\"https://www.kaggle.com/drzhuzhe/transformer\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/transformer</a></p>\n<p>Refferrences:</p>\n<ol>\n<li><p>It can work well with TPU , due to Darien Schettler's TPU strategies and dataset<br>\n<a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs</a></p></li>\n<li><p>Transformer model is build from Aditya Mishra's work<br>\n<a href=\"https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer/\" target=\"_blank\">https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer/</a></p></li>\n</ol>\n<p>I'm wonder which part I shall begin with ?<br>\ntransformer model architecture?  loss function? params ?</p>\n<p>Maybe someone can giåve me some suggestions like \"a self checklist for baseline model\" </p>\n<p>thanks , </p>\n<p>I'm keeping searching for tutorial or something , but not a solution yet</p>",
  "messages": [
    {
      "id": "1321907",
      "postDate": "05/25/2021 04:47:00",
      "content": "<p>This is my simple efficientnet-b2 + transformer model <br>\nit seems hardly handle with training loss<br>\n<a href=\"https://www.kaggle.com/drzhuzhe/transformer\" target=\"_blank\">https://www.kaggle.com/drzhuzhe/transformer</a></p>\n<p>Refferrences:</p>\n<ol>\n<li><p>It can work well with TPU , due to Darien Schettler's TPU strategies and dataset<br>\n<a href=\"https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\" target=\"_blank\">https://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs</a></p></li>\n<li><p>Transformer model is build from Aditya Mishra's work<br>\n<a href=\"https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer/\" target=\"_blank\">https://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer/</a></p></li>\n</ol>\n<p>I'm wonder which part I shall begin with ?<br>\ntransformer model architecture?  loss function? params ?</p>\n<p>Maybe someone can giåve me some suggestions like \"a self checklist for baseline model\" </p>\n<p>thanks , </p>\n<p>I'm keeping searching for tutorial or something , but not a solution yet</p>",
      "rawMarkdown": "This is my simple efficientnet-b2 + transformer model \nit seems hardly handle with training loss\nhttps://www.kaggle.com/drzhuzhe/transformer\n\nRefferrences:\n\n1. It can work well with TPU , due to Darien Schettler's TPU strategies and dataset\nhttps://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\n\n2. Transformer model is build from Aditya Mishra's work\nhttps://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer/\n\nI'm wonder which part I shall begin with ?\ntransformer model architecture?  loss function? params ?\n\nMaybe someone can giåve me some suggestions like \"a self checklist for baseline model\" \n\nthanks , \n\nI'm keeping searching for tutorial or something , but not a solution yet",
      "votes": null
    },
    {
      "id": "1322902",
      "postDate": "05/25/2021 19:20:34",
      "content": "<p>See here -&gt; <a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e?scriptVersionId=63942804\" target=\"_blank\">https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e?scriptVersionId=63942804</a></p>\n<p>I wrote a discussion post on it here -&gt; <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/241716\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/241716</a></p>",
      "rawMarkdown": "See here -> https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e?scriptVersionId=63942804\n\nI wrote a discussion post on it here -> https://www.kaggle.com/c/bms-molecular-translation/discussion/241716",
      "votes": null
    },
    {
      "id": "1323109",
      "postDate": "05/26/2021 02:19:03",
      "content": "<p>you are right spike learning rate is the key <br>\nin my first version model structure is already correct</p>",
      "rawMarkdown": "you are right spike learning rate is the key \nin my first version model structure is already correct",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1322902,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "05/25/2021 19:20:34",
      "content": "<p>See here -&gt; <a href=\"https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e?scriptVersionId=63942804\" target=\"_blank\">https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e?scriptVersionId=63942804</a></p>\n<p>I wrote a discussion post on it here -&gt; <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/241716\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/241716</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1323109,
          "author_name": "drzhuzhe",
          "author_url": "",
          "post_date": "05/26/2021 02:19:03",
          "content": "<p>you are right spike learning rate is the key <br>\nin my first version model structure is already correct</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1321907": "This is my simple efficientnet-b2 + transformer model \nit seems hardly handle with training loss\nhttps://www.kaggle.com/drzhuzhe/transformer\n\nRefferrences:\n\n1. It can work well with TPU , due to Darien Schettler's TPU strategies and dataset\nhttps://www.kaggle.com/dschettler8845/bms-efficientnetv2-tpu-e2e-pipeline-in-3hrs\n\n2. Transformer model is build from Aditya Mishra's work\nhttps://www.kaggle.com/aditya08/imagecaptioning-show-attnd-tell-w-transformer/\n\nI'm wonder which part I shall begin with ?\ntransformer model architecture?  loss function? params ?\n\nMaybe someone can giåve me some suggestions like \"a self checklist for baseline model\" \n\nthanks , \n\nI'm keeping searching for tutorial or something , but not a solution yet",
    "1322902": "See here -> https://www.kaggle.com/dschettler8845/3-71-cv-bms-efficientnetv2-transformer-e2e?scriptVersionId=63942804\n\nI wrote a discussion post on it here -> https://www.kaggle.com/c/bms-molecular-translation/discussion/241716",
    "1323109": "you are right spike learning rate is the key \nin my first version model structure is already correct"
  },
  "source": "meta"
}