{
  "id": 349715,
  "title": "Starting with Encoder-Decoder Neural Network",
  "url": "/competitions/open-problems-multimodal/discussion/349715",
  "author_name": "",
  "post_date": "2022-09-02T13:03:32.244338800Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/ravishah1/citeseq-rna-to-protein-encoder-decoder-nn\" target=\"_blank\">notebook / code example </a></p>\n<p>While I have been seeing a lot of linear and gbt based approaches, I thought it would be interesting to try making a neural network in pytorch. Hopefully this gives you a decent nn baseline.</p>\n<h3>Background</h3>\n<ul>\n<li>CITEseq samples take input X (RNA sequence vector) to predict output Y (Protein sequence vector)</li>\n<li>Import generated features (n_samples, 240) shape created using PCA and simple feature engineering</li>\n<li>Feed samples to encoder-decoder NN (see structure below)</li>\n<li>Train one fold using pytorch (20 epochs, AdamW optimizer, and Cosine scheduler)</li>\n<li>Use network to make predictions</li>\n</ul>\n<h3>Improving Upon this Notebook</h3>\n<p><strong>The best way to improve upon this notebook is likely feature engineering</strong></p>\n<ul>\n<li>Look for feature importance</li>\n<li>Improve dimensionality reduction technique</li>\n<li>Use domain knowledge</li>\n<li>This is just a baseline</li>\n</ul>\n<p><strong>Longer Training</strong></p>\n<ul>\n<li>This notebook is only trained for 50 epochs (only 40 min) and improved every epoch</li>\n<li>This notebook only shows the results from one fold, you can train all folds</li>\n</ul>\n<p><strong>Apply similar model for Multiome samples</strong> </p>\n<ul>\n<li>Currently, I am just borrowing another submission for multiome and this notebook only predicts CITEseq</li>\n<li>However, you can expand upon this notebook to make Multiome predictions</li>\n</ul>\n<p><strong>Change NN Structure</strong></p>\n<ul>\n<li>test rnn or cnn</li>\n<li>try adding attention mechanism</li>\n<li>change structure to adjust for new features</li>\n</ul>\n<h3>Neural Network Structure</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcc192e980bc9f9248d2bcae2c6accd74%2Fencoder_decoder.PNG?generation=1662083041543650&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1a17ab66143625efff11e8a063e1dac1%2Fenc_dec2.PNG?generation=1662083054477703&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb373fe534194a31dfc3505f90f488b25%2FFCBlock.PNG?generation=1662087114666150&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "1923678",
      "postDate": "09/02/2022 13:03:32",
      "content": "<p><a href=\"https://www.kaggle.com/code/ravishah1/citeseq-rna-to-protein-encoder-decoder-nn\" target=\"_blank\">notebook / code example </a></p>\n<p>While I have been seeing a lot of linear and gbt based approaches, I thought it would be interesting to try making a neural network in pytorch. Hopefully this gives you a decent nn baseline.</p>\n<h3>Background</h3>\n<ul>\n<li>CITEseq samples take input X (RNA sequence vector) to predict output Y (Protein sequence vector)</li>\n<li>Import generated features (n_samples, 240) shape created using PCA and simple feature engineering</li>\n<li>Feed samples to encoder-decoder NN (see structure below)</li>\n<li>Train one fold using pytorch (20 epochs, AdamW optimizer, and Cosine scheduler)</li>\n<li>Use network to make predictions</li>\n</ul>\n<h3>Improving Upon this Notebook</h3>\n<p><strong>The best way to improve upon this notebook is likely feature engineering</strong></p>\n<ul>\n<li>Look for feature importance</li>\n<li>Improve dimensionality reduction technique</li>\n<li>Use domain knowledge</li>\n<li>This is just a baseline</li>\n</ul>\n<p><strong>Longer Training</strong></p>\n<ul>\n<li>This notebook is only trained for 50 epochs (only 40 min) and improved every epoch</li>\n<li>This notebook only shows the results from one fold, you can train all folds</li>\n</ul>\n<p><strong>Apply similar model for Multiome samples</strong> </p>\n<ul>\n<li>Currently, I am just borrowing another submission for multiome and this notebook only predicts CITEseq</li>\n<li>However, you can expand upon this notebook to make Multiome predictions</li>\n</ul>\n<p><strong>Change NN Structure</strong></p>\n<ul>\n<li>test rnn or cnn</li>\n<li>try adding attention mechanism</li>\n<li>change structure to adjust for new features</li>\n</ul>\n<h3>Neural Network Structure</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcc192e980bc9f9248d2bcae2c6accd74%2Fencoder_decoder.PNG?generation=1662083041543650&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1a17ab66143625efff11e8a063e1dac1%2Fenc_dec2.PNG?generation=1662083054477703&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb373fe534194a31dfc3505f90f488b25%2FFCBlock.PNG?generation=1662087114666150&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "[notebook / code example ](https://www.kaggle.com/code/ravishah1/citeseq-rna-to-protein-encoder-decoder-nn)\n\nWhile I have been seeing a lot of linear and gbt based approaches, I thought it would be interesting to try making a neural network in pytorch. Hopefully this gives you a decent nn baseline.\n\n### Background\n\n- CITEseq samples take input X (RNA sequence vector) to predict output Y (Protein sequence vector)\n- Import generated features (n_samples, 240) shape created using PCA and simple feature engineering\n- Feed samples to encoder-decoder NN (see structure below)\n- Train one fold using pytorch (20 epochs, AdamW optimizer, and Cosine scheduler)\n- Use network to make predictions\n\n### Improving Upon this Notebook\n\n**The best way to improve upon this notebook is likely feature engineering**\n- Look for feature importance\n- Improve dimensionality reduction technique\n- Use domain knowledge\n- This is just a baseline\n\n**Longer Training**\n- This notebook is only trained for 50 epochs (only 40 min) and improved every epoch\n- This notebook only shows the results from one fold, you can train all folds\n\n**Apply similar model for Multiome samples** \n- Currently, I am just borrowing another submission for multiome and this notebook only predicts CITEseq\n- However, you can expand upon this notebook to make Multiome predictions\n\n**Change NN Structure**\n- test rnn or cnn\n- try adding attention mechanism\n- change structure to adjust for new features\n\n### Neural Network Structure\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcc192e980bc9f9248d2bcae2c6accd74%2Fencoder_decoder.PNG?generation=1662083041543650&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1a17ab66143625efff11e8a063e1dac1%2Fenc_dec2.PNG?generation=1662083054477703&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb373fe534194a31dfc3505f90f488b25%2FFCBlock.PNG?generation=1662087114666150&alt=media)",
      "votes": null
    },
    {
      "id": "1924489",
      "postDate": "09/03/2022 05:24:35",
      "content": "<p>expect result of transformer</p>",
      "rawMarkdown": "expect result of transformer",
      "votes": null
    },
    {
      "id": "1947874",
      "postDate": "09/20/2022 17:28:24",
      "content": "<p>how is this performing after the rescore?</p>",
      "rawMarkdown": "how is this performing after the rescore?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1924489,
      "author_name": "huangzchao",
      "author_url": "",
      "post_date": "09/03/2022 05:24:35",
      "content": "<p>expect result of transformer</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1947874,
      "author_name": "mohami",
      "author_url": "",
      "post_date": "09/20/2022 17:28:24",
      "content": "<p>how is this performing after the rescore?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1923678": "[notebook / code example ](https://www.kaggle.com/code/ravishah1/citeseq-rna-to-protein-encoder-decoder-nn)\n\nWhile I have been seeing a lot of linear and gbt based approaches, I thought it would be interesting to try making a neural network in pytorch. Hopefully this gives you a decent nn baseline.\n\n### Background\n\n- CITEseq samples take input X (RNA sequence vector) to predict output Y (Protein sequence vector)\n- Import generated features (n_samples, 240) shape created using PCA and simple feature engineering\n- Feed samples to encoder-decoder NN (see structure below)\n- Train one fold using pytorch (20 epochs, AdamW optimizer, and Cosine scheduler)\n- Use network to make predictions\n\n### Improving Upon this Notebook\n\n**The best way to improve upon this notebook is likely feature engineering**\n- Look for feature importance\n- Improve dimensionality reduction technique\n- Use domain knowledge\n- This is just a baseline\n\n**Longer Training**\n- This notebook is only trained for 50 epochs (only 40 min) and improved every epoch\n- This notebook only shows the results from one fold, you can train all folds\n\n**Apply similar model for Multiome samples** \n- Currently, I am just borrowing another submission for multiome and this notebook only predicts CITEseq\n- However, you can expand upon this notebook to make Multiome predictions\n\n**Change NN Structure**\n- test rnn or cnn\n- try adding attention mechanism\n- change structure to adjust for new features\n\n### Neural Network Structure\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcc192e980bc9f9248d2bcae2c6accd74%2Fencoder_decoder.PNG?generation=1662083041543650&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F1a17ab66143625efff11e8a063e1dac1%2Fenc_dec2.PNG?generation=1662083054477703&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb373fe534194a31dfc3505f90f488b25%2FFCBlock.PNG?generation=1662087114666150&alt=media)",
    "1924489": "expect result of transformer",
    "1947874": "how is this performing after the rescore?"
  },
  "source": "meta"
}