{
  "id": 461649,
  "title": "#30: \"Melting\" the Data",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/461649",
  "author_name": "Frenio Redeker",
  "post_date": "2023-12-15T16:29:44.682000",
  "votes": 8,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Thank you to the organizers of the Open Problems – Single-Cell Perturbations competition for this incredible opportunity. </p>\n<p>My approach combines four model types: Tabular Models with embeddings and dense layers, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests, all trained on a \"melted\" format of the training data. </p>\n<p>In the following sections, I provide a comprehensive breakdown of my approach and the specific models employed. Links to models and other resources are provided throughout the text and can also be found in Section 6.</p>\n<p></p>\n<h2>1. Integration of biological knowledge</h2>\n<p>My initial approach involved training Tabular Neural Networks and Transformers specifically on the categorical features available in the training data. These models excel in encoding complex relationships within their embedding layers. Concurrently, I planned to augment the training data with information sourced from biological databases and web searches, and I intended to incorporate calculated molecular descriptors derived from SMILES. This enriched dataset would then be used to train tree-based models.</p>\n<h3>Molecular Descriptors</h3>\n<p>I calculated 1600+ molecular descriptors using the <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-018-0258-y\" target=\"_blank\">mordred</a> python library and trained a random forest using all molecular descriptors in order to select the most important features (importance &gt; 0.5 %) based on the feature importance results of the training run, which yielded a set of 23 molecular descriptors. All molecular descriptors as well as the list of most important descriptors can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a> in the files <code>mol_descriptors.parquet</code> and <code>important_mol_descriptors.csv</code>, respectively.</p>\n<pre><code> rdkit, rdkit.Chem\n mordred, mordred.descriptors\n\ncalc = mordred.Calculator(mordred.descriptors, ignore_3D=)\nmolecules = [rdkit.Chem.MolFromSmiles(smi)  smi  compounds[]]\nfeatures = calc.pandas(molecules)\n</code></pre>\n<p>Here <code>compounds['SMILES']</code> contained a list of unique SMILES strings from the training data.</p>\n<h3>Gene Information</h3>\n<p>Gene information was obtained from the ENSEMBL database through the <a href=\"https://jrderuiter.github.io/pybiomart/index.html\" target=\"_blank\">pybiomart</a> python package. Note that gene info could only be obtained for 13,423 of the 18,211 genes in this way.</p>\n<pre><code> pybiomart  Server\n\nserver = Server(host=)\ndataset = server[][]\nresults = dataset.query(attributes=[, , , , ], filters={: id_filter, : })\n</code></pre>\n<p>From there I added additional features. For example, the number of each nucleobase (ACGT) obtained from the sequence data seemed particularly useful, according to feature importance results evaluated after training of Random Forests. A data frame containing added gene info can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a> as <code>geneinfo.parquet</code>.</p>\n<h3>Cell Features</h3>\n<p>Finding useful features that distinguish the 6 different cell types in the training data was a challenge for me. I ended up browsing Wikipedia, PubMed, and the <a href=\"https://www.immunology.org/public-information/bitesized-immunology/cells\" target=\"_blank\">website</a> of the British Society for Immunology for information that I could add as cell features. The resulting data frame can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a> as <code>cellinfo.csv</code>.</p>\n<h3>Result of Direct Integration of Biological Knowledge</h3>\n<p>The best result that I was able to obtain using a tree-based model after adding the features described above was a Random Forest that achieved scores of 0.628 and 0.836 on the public and private leaderboards, respectively. For comparison, my best Random Forest trained on embeddings learned by a Tabular Neural Network (TabMod NN) achieved scores of 0.588 and 0.780.</p>\n<h3>Embeddings</h3>\n<p>Due to the inadequacy of direct integration of molecular descriptors, gene, and cell information, I decided to train Random and XGBoosted Forests on embeddings learned by the TabMod Neural Network. The use of embeddings significantly improved my results and might reasonably be considered an indirect form of \"integration of biological knowledge,\" because neural networks like TabMod NN are able to uncover and encode complex, often non-linear relationships inherent in biological systems into these embedding vectors. The embeddings obtained from TabMod NN can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings\" target=\"_blank\">OP2 TabMod NN Embeddings Dataset</a>.</p>\n<p>Plots of the <code>cell_type</code>, <code>sm_name</code>, and <code>gene</code> embeddings after PCA using 2 components as well as the corresponding code can be found in the <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings\" target=\"_blank\">Look at Embeddings Notebook</a>. The plot of the <code>gene</code> embeddings shows two large distinct clusters – one very dense, the other less dense. The red marks with labels show a random selection of gene names to be displayed.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2Fa3a4ee89d9b9f0bc7d51f7c55e6faad6%2Fgene_embeds.png?generation=1702353451391406&amp;alt=media\" alt=\"Gene Embeddings\"></p>\n<h2>2. Exploration of the problem</h2>\n<p>At first glance, the training data seemed like a collaborative filtering problem, where cell types are the users, drugs are the items, and gene expression confidence values are the targets. However, thinking about it that way leads to 18211 collaborative filtering problems – one for every gene – and I was intimidated by the idea of training 18211 models to come up with a single submission. Then I realized that collaborative filtering is just a special case of a tabular modeling problem with two features and one target, which led to the idea to \"melt\" the training data and the test data (convert from wide to long format):</p>\n<pre><code>train_df = df_de_train.melt(id_vars=[, ], value_vars=df_de_train.iloc[:,:].columns, var_name=, value_name=)\n</code></pre>\n<p>Melting of the training data results in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (<code>cell_type</code>, <code>sm_name</code> and <code>gene</code>), and one target (signed -log(p-value) called <code>value</code>). I used this data format for all my models, which then had to predict only one target for each <code>cell_type</code>/<code>sm_name</code>/<code>gene</code> combination. Inference using a model trained on data in that format yielded a list of 4,643,805 prediction values, which I just reshaped back into the submission format of 255 x 18,211 (see e.g. my <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds\" target=\"_blank\">Random Forest Notebook</a>).</p>\n<h2>3. Model design</h2>\n<p>For final submission, I used an ensemble of 4 different model types: Tabular Model Neural Networks, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests. These are described in more detail below. The ensemble that achieved the placement in position 30 of the private leader board (with scores of 0.553 and 0.753 in the public and private leaderboards, respectively) had the following structure:</p>\n<p>5 TabMod NNs using 600 dimensions for gene embeddings and random seeds 42, 55, 120, 457, and 736 and 5 TabMod NNs using 1000 dimensions for gene embeddings and random seeds 42, 199, 550, 855, and 970:</p>\n<pre><code>df1 = df1_1* + df1_2* + df1_3* + df1_4* + df1_5* + df1_6* + df1_7* + df1_8* + df1_9* + df1_10*\n</code></pre>\n<p>5 Random Forests using random seeds 209, 569, 739, 885, and 926:</p>\n<pre><code>df2 = df2_1* + df2_2* + df2_3* + df2_4* + df2_5*\n</code></pre>\n<p>5 XGBoosted Forests using random seeds 117, 150, 234, 624, and 804:</p>\n<pre><code>df3 = df3_1* + df3_2* + df3_3* + df3_4* + df3_5*\n</code></pre>\n<p>5 fine-tuned Transformer Models, TinyBioBert and DeBERTa-v3-small (see details below):</p>\n<pre><code>df4 = df4_1* + df4_2* + df4_3* + df4_4* + df4_5*\n</code></pre>\n<p>The weights used for the final submission were:</p>\n<pre><code>submission = df1* + df2* + df3* + df4*\n</code></pre>\n<p>The following figure shows a correlation heat map of the predictions of the different model types (see the <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions\" target=\"_blank\">Prediction Correlation Notebook</a> for the source code):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F261951234c01103ce2cb74c09c7dad41%2Fprediction_correlation.png?generation=1702353628253144&amp;alt=media\" alt=\"Prediction Correlation\"></p>\n<p>In the subsequent sections I present detailed descriptions of each model type, along with links to the relevant Kaggle notebooks. These notebooks are designed to run on Kaggle. However, the fine-tuning of transformer models, utilizing the full training data, was conducted on SaturnCloud. To adapt these notebooks for Kaggle, the training data was substantially reduced. As a result, the transformer models showcased on Kaggle are for demonstrative purposes only and do not achieve scores indicative of their full potential in the competition.</p>\n<h3>TabMod NN</h3>\n<p>TabMod NN performance was best when the training data was denoised using PCA and 10 components prior to training (see the <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising\" target=\"_blank\">TabMod NN Notebook</a> for details and code). The two kinds of tabular neural networks used for final submission were identical except for the dimensionality (600 and 1000) of gene embeddings used. The models were based on fast.ai’s <code>tabular_learner</code> with the following configuration:</p>\n<pre><code>learn = tabular_learner(dls, y_range=(y_min, y_max), emb_szs=emb_szs, layers=[, , ], n_out=, loss_func=F.mse_loss)\n</code></pre>\n<p>To achieve a gene embedding size of 1000 instead of the default 389, the <code>emb_szs</code> dictionary was customized.</p>\n<pre><code>emb_szs = {: , : , : }\n</code></pre>\n<p>The range of the final sigmoid layer was set to <code>y_min</code> to <code>y_max</code> which were obtained by determining the minimum and maximum target values in the training data after PCA denoising. The three dense layers of size 1000, 500, and 250 were optimized for appropriate expressivity. Finally, the mean squared loss function was chosen to best reflect the evaluation in the leaderboard. Experiments using mean absolute error, huber loss, and log-cosh did not lead to improvements of the model and instead reduced model performance.</p>\n<h3>Random Forest</h3>\n<p>Random Forests (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds\" target=\"_blank\">Random Forest Notebook</a>) were trained after adding embeddings for <code>cell_type</code>, <code>sm_name</code>, and <code>gene</code> to the training data. These embeddings were obtained from a TabMod NN similar to the one described in the previous section with gene embeddings of dimensionality 1000 but without the use of PCA denoising on the training data (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-export-tabmod-nn-embeddings\" target=\"_blank\">Export Embeddings Notebook</a> for the code). Of those the <code>cell_type</code> embeddings of dimensionality 5 were left unchanged, whereas <code>sm_name</code> and <code>gene</code> embeddings (of dimensionality 26 and 1000, respectively) were reduced to 10 components each using PCA.</p>\n<p>All Random Forest models used in the final submission were identical except for the random seeds used. They were based on scikitlearn’s <code>RandomForestRegressor</code> using 100 trees on 66 % of samples in order to achieve a meaningful out-of-bag error and score. The square root of feature number was used as the maximum number of features per tree which has been <a href=\"https://scikit-learn.org/stable/auto_examples/ensemble/plot_ensemble_oob.html\" target=\"_blank\">shown</a> to be beneficial for the model’s ability to generalize. The minimum number of samples per leaf was set to 5 in order to achieve high expressivity. Such a Random Forest could be trained using the following example code:</p>\n<pre><code> sklearn.ensemble  RandomForestRegressor\n\nm = RandomForestRegressor(n_jobs=-, n_estimators=, max_samples=, max_features=, min_samples_leaf=, oob_score=)\nm.fit(xs, y)\n</code></pre>\n<h3>XGBoosted Forest</h3>\n<p>The same embeddings used to train Random Forests were also used to train XGBoosted Forests using the same dimensionality reductions.</p>\n<p>The XGBoosted Forests used in the final submission were also identical to each other except for the random seed used. They were trained using the following code (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds\" target=\"_blank\">XGBoosted Forest Notebook</a> for details):</p>\n<pre><code> xgboost  xgb\n\nm = xgb.XGBRegressor(device=, n_estimators=, learning_rate=, max_depth=, min_child_weight=, gamma=, subsample=, colsample_bytree=)\nm.fit(xs, y)\n</code></pre>\n<p>The model parameters were optimized using the cross validation scheme proposed by <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&amp;cellId=8\" target=\"_blank\">AmbrosM</a>, where all but 10 % of the data of one of four cell types is used as a validation set. It should be noted that the performance of the model could potentially be enhanced by configuring the <code>max_depth</code> parameter to values exceeding 20. This approach was not extensively explored due to the higher computational costs associated with larger values.</p>\n<h3>Transformer Models</h3>\n<p>Two kinds of Transformer models from the <a href=\"https://huggingface.co/docs/transformers/index\" target=\"_blank\">Huggingface Transformers</a> data base were fine-tuned to output a number when given a standardized input sentence. The general model configuration is shown in the following code snippet:</p>\n<pre><code> transformers  TrainingArguments, Trainer\n\nargs = TrainingArguments(, save_steps=steps, learning_rate=, warmup_ratio=, lr_scheduler_type=, fp16=,\n    evaluation_strategy=, per_device_train_batch_size=bs, per_device_eval_batch_size=bs*,\n    num_train_epochs=epochs, weight_decay=, report_to=, seed=random_seed)\n\nmodel = AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=)\n</code></pre>\n<p>The argument <code>num_labels=1</code> configures the final linear layer of the model to have a single output neuron which is appropriate for a regression task, as it leads to prediction of a single continuous value.</p>\n<p>The first model was based on the <a href=\"https://huggingface.co/microsoft/deberta-v3-small\" target=\"_blank\">deberta-v3-small</a> model (<a href=\"https://openreview.net/forum?id=XPZIaotutsD\" target=\"_blank\">link to article</a>) with 44M parameters pre-trained on general text data. The \"sparse\" input used for the DeBERTa model was:</p>\n<pre><code>trdf[] =  + trdf.cell_type +  + trdf.sm_name +  + trdf.gene\n</code></pre>\n<p>Fine-tuning the DeBERTa for 20-25 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 48-60 hours (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo\" target=\"_blank\">Transformer DeBERTa (Demo) Notebook</a> for a training demo that runs on Kaggle in ~2 hours).</p>\n<p>The second model was based on the <a href=\"https://huggingface.co/nlpie/tiny-biobert\" target=\"_blank\">tiny-biobert</a> model (<a href=\"https://doi.org/10.48550/arxiv.2209.03182\" target=\"_blank\">link to article</a>) with 15M parameters pre-trained on citations and abstracts of biomedical literature using the PubMed dataset. The \"verbose\" input used for the TinyBioBert model was:</p>\n<pre><code>trdf[] =  + trdf.gene +  + trdf.cell_type +  + trdf.sm_name + \n</code></pre>\n<p>Fine-tuning the TinyBioBert model for 20-30 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 24-30 hours (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo\" target=\"_blank\">Transformer TinyBioBert (Demo) Notebook</a> for a training demo that runs on Kaggle in ~1:30 hours).</p>\n<p>Interestingly, the \"sparse\" input used for the DeBERTa model did not lead to very good performance in the TinyBioBert model, whereas the more \"verbose\" input used for the TinyBioBert model led to slower training and worse results in the DeBERTa model. This is reflected in the following submission scores (all trained for 5 epochs using the same random seed of 42):</p>\n<pre><code>                        public LB     privateLB\nDeBERTa sparse:                      \nDeBERTa verbose:                     \nTinyBioBert sparse:                  \nTinyBioBert verbose:                 \n</code></pre>\n<p>The transformer part of the final ensemble consisted of four TinyBioBert models. These were trained for 30 epochs using seed 42, 30 epochs using seed 546, 20 epochs using seed 419, and 5x5 epochs using different random seeds. Each model was assigned a weight of 0.15 in the Transformer sub-ensemble. The single DeBERTa model used was trained for a total of 22 epochs using different random seeds, and was assigned a weight of 0.4.</p>\n<h2>4. Robustness</h2>\n<p>The robustness of each individual model type was evaluated using triplicate submissions with different random seeds by calculating the mean and standard deviation of the scores achieved on the public and private leaderboard (in the case of transformers, the models also varied slightly in the number of epochs trained). </p>\n<p>TabMod NN (600): 0.567 ± 0.003 public, 0.777 ± 0.007 private.<br>\nTabMod NN (1000): 0.562 ± 0.002 public, 0.778 ± 0.003 private</p>\n<p>Random Forest: 0.592 ± 0.004 public, 0.781 ± 0.002 private.</p>\n<p>XGBoosted Forest: 0.588 ± 0.002 public, 0.780 ± 0.001 private.</p>\n<p>Transformer (DeBERTa) 22-25 epochs: 0.606 ± 0.013 public, 0.784 ± 0.007 private.<br>\nTransformer (TinyBioBert) 25-30 epochs: 0.607 ± 0.003 public, 0.777 ± 0.004 private.</p>\n<p>These results show that the fine-tuned DeBERTa model is the least robust of the models. The fine-tuned TinyBioBert model consistently performed as well as the TabMod NNs on the private leaderboard, but consistently looked worse on the public leaderboard.</p>\n<h2>5. Documentation &amp; code style</h2>\n<p>The code is documented in the notebooks linked throughout the notebook. Additionally, a list of links to all notebooks and datasets can be found in Section 6.</p>\n<h2>6. Reproducibility</h2>\n<p>The source code and further documentation for all models as well as additional datasets can be accessed through the following links:</p>\n<p>Model Notebooks:<br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising\" target=\"_blank\">TabMod NN Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds\" target=\"_blank\">Random Forest Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds\" target=\"_blank\">XGBoosted Forest Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo\" target=\"_blank\">Transformer DeBERTa (Demo) Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo\" target=\"_blank\">Transformer TinyBioBert (Demo) Notebook</a></p>\n<p>Other Notebooks:<br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings\" target=\"_blank\">Look at Embeddings Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions\" target=\"_blank\">Prediction Correlation Notebook</a></p>\n<p>Datasets:<br>\n<a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a><br>\n<a href=\"https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings\" target=\"_blank\">OP2 TabMod NN Embeddings Dataset</a></p>",
  "messages": [
    {
      "id": 2562753,
      "postDate": "2023-12-15T16:29:44.683Z",
      "content": "<p>Thank you to the organizers of the Open Problems – Single-Cell Perturbations competition for this incredible opportunity. </p>\n<p>My approach combines four model types: Tabular Models with embeddings and dense layers, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests, all trained on a \"melted\" format of the training data. </p>\n<p>In the following sections, I provide a comprehensive breakdown of my approach and the specific models employed. Links to models and other resources are provided throughout the text and can also be found in Section 6.</p>\n<p></p>\n<h2>1. Integration of biological knowledge</h2>\n<p>My initial approach involved training Tabular Neural Networks and Transformers specifically on the categorical features available in the training data. These models excel in encoding complex relationships within their embedding layers. Concurrently, I planned to augment the training data with information sourced from biological databases and web searches, and I intended to incorporate calculated molecular descriptors derived from SMILES. This enriched dataset would then be used to train tree-based models.</p>\n<h3>Molecular Descriptors</h3>\n<p>I calculated 1600+ molecular descriptors using the <a href=\"https://jcheminf.biomedcentral.com/articles/10.1186/s13321-018-0258-y\" target=\"_blank\">mordred</a> python library and trained a random forest using all molecular descriptors in order to select the most important features (importance &gt; 0.5 %) based on the feature importance results of the training run, which yielded a set of 23 molecular descriptors. All molecular descriptors as well as the list of most important descriptors can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a> in the files <code>mol_descriptors.parquet</code> and <code>important_mol_descriptors.csv</code>, respectively.</p>\n<pre><code> rdkit, rdkit.Chem\n mordred, mordred.descriptors\n\ncalc = mordred.Calculator(mordred.descriptors, ignore_3D=)\nmolecules = [rdkit.Chem.MolFromSmiles(smi)  smi  compounds[]]\nfeatures = calc.pandas(molecules)\n</code></pre>\n<p>Here <code>compounds['SMILES']</code> contained a list of unique SMILES strings from the training data.</p>\n<h3>Gene Information</h3>\n<p>Gene information was obtained from the ENSEMBL database through the <a href=\"https://jrderuiter.github.io/pybiomart/index.html\" target=\"_blank\">pybiomart</a> python package. Note that gene info could only be obtained for 13,423 of the 18,211 genes in this way.</p>\n<pre><code> pybiomart  Server\n\nserver = Server(host=)\ndataset = server[][]\nresults = dataset.query(attributes=[, , , , ], filters={: id_filter, : })\n</code></pre>\n<p>From there I added additional features. For example, the number of each nucleobase (ACGT) obtained from the sequence data seemed particularly useful, according to feature importance results evaluated after training of Random Forests. A data frame containing added gene info can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a> as <code>geneinfo.parquet</code>.</p>\n<h3>Cell Features</h3>\n<p>Finding useful features that distinguish the 6 different cell types in the training data was a challenge for me. I ended up browsing Wikipedia, PubMed, and the <a href=\"https://www.immunology.org/public-information/bitesized-immunology/cells\" target=\"_blank\">website</a> of the British Society for Immunology for information that I could add as cell features. The resulting data frame can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a> as <code>cellinfo.csv</code>.</p>\n<h3>Result of Direct Integration of Biological Knowledge</h3>\n<p>The best result that I was able to obtain using a tree-based model after adding the features described above was a Random Forest that achieved scores of 0.628 and 0.836 on the public and private leaderboards, respectively. For comparison, my best Random Forest trained on embeddings learned by a Tabular Neural Network (TabMod NN) achieved scores of 0.588 and 0.780.</p>\n<h3>Embeddings</h3>\n<p>Due to the inadequacy of direct integration of molecular descriptors, gene, and cell information, I decided to train Random and XGBoosted Forests on embeddings learned by the TabMod Neural Network. The use of embeddings significantly improved my results and might reasonably be considered an indirect form of \"integration of biological knowledge,\" because neural networks like TabMod NN are able to uncover and encode complex, often non-linear relationships inherent in biological systems into these embedding vectors. The embeddings obtained from TabMod NN can be found in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings\" target=\"_blank\">OP2 TabMod NN Embeddings Dataset</a>.</p>\n<p>Plots of the <code>cell_type</code>, <code>sm_name</code>, and <code>gene</code> embeddings after PCA using 2 components as well as the corresponding code can be found in the <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings\" target=\"_blank\">Look at Embeddings Notebook</a>. The plot of the <code>gene</code> embeddings shows two large distinct clusters – one very dense, the other less dense. The red marks with labels show a random selection of gene names to be displayed.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2Fa3a4ee89d9b9f0bc7d51f7c55e6faad6%2Fgene_embeds.png?generation=1702353451391406&amp;alt=media\" alt=\"Gene Embeddings\"></p>\n<h2>2. Exploration of the problem</h2>\n<p>At first glance, the training data seemed like a collaborative filtering problem, where cell types are the users, drugs are the items, and gene expression confidence values are the targets. However, thinking about it that way leads to 18211 collaborative filtering problems – one for every gene – and I was intimidated by the idea of training 18211 models to come up with a single submission. Then I realized that collaborative filtering is just a special case of a tabular modeling problem with two features and one target, which led to the idea to \"melt\" the training data and the test data (convert from wide to long format):</p>\n<pre><code>train_df = df_de_train.melt(id_vars=[, ], value_vars=df_de_train.iloc[:,:].columns, var_name=, value_name=)\n</code></pre>\n<p>Melting of the training data results in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (<code>cell_type</code>, <code>sm_name</code> and <code>gene</code>), and one target (signed -log(p-value) called <code>value</code>). I used this data format for all my models, which then had to predict only one target for each <code>cell_type</code>/<code>sm_name</code>/<code>gene</code> combination. Inference using a model trained on data in that format yielded a list of 4,643,805 prediction values, which I just reshaped back into the submission format of 255 x 18,211 (see e.g. my <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds\" target=\"_blank\">Random Forest Notebook</a>).</p>\n<h2>3. Model design</h2>\n<p>For final submission, I used an ensemble of 4 different model types: Tabular Model Neural Networks, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests. These are described in more detail below. The ensemble that achieved the placement in position 30 of the private leader board (with scores of 0.553 and 0.753 in the public and private leaderboards, respectively) had the following structure:</p>\n<p>5 TabMod NNs using 600 dimensions for gene embeddings and random seeds 42, 55, 120, 457, and 736 and 5 TabMod NNs using 1000 dimensions for gene embeddings and random seeds 42, 199, 550, 855, and 970:</p>\n<pre><code>df1 = df1_1* + df1_2* + df1_3* + df1_4* + df1_5* + df1_6* + df1_7* + df1_8* + df1_9* + df1_10*\n</code></pre>\n<p>5 Random Forests using random seeds 209, 569, 739, 885, and 926:</p>\n<pre><code>df2 = df2_1* + df2_2* + df2_3* + df2_4* + df2_5*\n</code></pre>\n<p>5 XGBoosted Forests using random seeds 117, 150, 234, 624, and 804:</p>\n<pre><code>df3 = df3_1* + df3_2* + df3_3* + df3_4* + df3_5*\n</code></pre>\n<p>5 fine-tuned Transformer Models, TinyBioBert and DeBERTa-v3-small (see details below):</p>\n<pre><code>df4 = df4_1* + df4_2* + df4_3* + df4_4* + df4_5*\n</code></pre>\n<p>The weights used for the final submission were:</p>\n<pre><code>submission = df1* + df2* + df3* + df4*\n</code></pre>\n<p>The following figure shows a correlation heat map of the predictions of the different model types (see the <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions\" target=\"_blank\">Prediction Correlation Notebook</a> for the source code):</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F261951234c01103ce2cb74c09c7dad41%2Fprediction_correlation.png?generation=1702353628253144&amp;alt=media\" alt=\"Prediction Correlation\"></p>\n<p>In the subsequent sections I present detailed descriptions of each model type, along with links to the relevant Kaggle notebooks. These notebooks are designed to run on Kaggle. However, the fine-tuning of transformer models, utilizing the full training data, was conducted on SaturnCloud. To adapt these notebooks for Kaggle, the training data was substantially reduced. As a result, the transformer models showcased on Kaggle are for demonstrative purposes only and do not achieve scores indicative of their full potential in the competition.</p>\n<h3>TabMod NN</h3>\n<p>TabMod NN performance was best when the training data was denoised using PCA and 10 components prior to training (see the <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising\" target=\"_blank\">TabMod NN Notebook</a> for details and code). The two kinds of tabular neural networks used for final submission were identical except for the dimensionality (600 and 1000) of gene embeddings used. The models were based on fast.ai’s <code>tabular_learner</code> with the following configuration:</p>\n<pre><code>learn = tabular_learner(dls, y_range=(y_min, y_max), emb_szs=emb_szs, layers=[, , ], n_out=, loss_func=F.mse_loss)\n</code></pre>\n<p>To achieve a gene embedding size of 1000 instead of the default 389, the <code>emb_szs</code> dictionary was customized.</p>\n<pre><code>emb_szs = {: , : , : }\n</code></pre>\n<p>The range of the final sigmoid layer was set to <code>y_min</code> to <code>y_max</code> which were obtained by determining the minimum and maximum target values in the training data after PCA denoising. The three dense layers of size 1000, 500, and 250 were optimized for appropriate expressivity. Finally, the mean squared loss function was chosen to best reflect the evaluation in the leaderboard. Experiments using mean absolute error, huber loss, and log-cosh did not lead to improvements of the model and instead reduced model performance.</p>\n<h3>Random Forest</h3>\n<p>Random Forests (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds\" target=\"_blank\">Random Forest Notebook</a>) were trained after adding embeddings for <code>cell_type</code>, <code>sm_name</code>, and <code>gene</code> to the training data. These embeddings were obtained from a TabMod NN similar to the one described in the previous section with gene embeddings of dimensionality 1000 but without the use of PCA denoising on the training data (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-export-tabmod-nn-embeddings\" target=\"_blank\">Export Embeddings Notebook</a> for the code). Of those the <code>cell_type</code> embeddings of dimensionality 5 were left unchanged, whereas <code>sm_name</code> and <code>gene</code> embeddings (of dimensionality 26 and 1000, respectively) were reduced to 10 components each using PCA.</p>\n<p>All Random Forest models used in the final submission were identical except for the random seeds used. They were based on scikitlearn’s <code>RandomForestRegressor</code> using 100 trees on 66 % of samples in order to achieve a meaningful out-of-bag error and score. The square root of feature number was used as the maximum number of features per tree which has been <a href=\"https://scikit-learn.org/stable/auto_examples/ensemble/plot_ensemble_oob.html\" target=\"_blank\">shown</a> to be beneficial for the model’s ability to generalize. The minimum number of samples per leaf was set to 5 in order to achieve high expressivity. Such a Random Forest could be trained using the following example code:</p>\n<pre><code> sklearn.ensemble  RandomForestRegressor\n\nm = RandomForestRegressor(n_jobs=-, n_estimators=, max_samples=, max_features=, min_samples_leaf=, oob_score=)\nm.fit(xs, y)\n</code></pre>\n<h3>XGBoosted Forest</h3>\n<p>The same embeddings used to train Random Forests were also used to train XGBoosted Forests using the same dimensionality reductions.</p>\n<p>The XGBoosted Forests used in the final submission were also identical to each other except for the random seed used. They were trained using the following code (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds\" target=\"_blank\">XGBoosted Forest Notebook</a> for details):</p>\n<pre><code> xgboost  xgb\n\nm = xgb.XGBRegressor(device=, n_estimators=, learning_rate=, max_depth=, min_child_weight=, gamma=, subsample=, colsample_bytree=)\nm.fit(xs, y)\n</code></pre>\n<p>The model parameters were optimized using the cross validation scheme proposed by <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&amp;cellId=8\" target=\"_blank\">AmbrosM</a>, where all but 10 % of the data of one of four cell types is used as a validation set. It should be noted that the performance of the model could potentially be enhanced by configuring the <code>max_depth</code> parameter to values exceeding 20. This approach was not extensively explored due to the higher computational costs associated with larger values.</p>\n<h3>Transformer Models</h3>\n<p>Two kinds of Transformer models from the <a href=\"https://huggingface.co/docs/transformers/index\" target=\"_blank\">Huggingface Transformers</a> data base were fine-tuned to output a number when given a standardized input sentence. The general model configuration is shown in the following code snippet:</p>\n<pre><code> transformers  TrainingArguments, Trainer\n\nargs = TrainingArguments(, save_steps=steps, learning_rate=, warmup_ratio=, lr_scheduler_type=, fp16=,\n    evaluation_strategy=, per_device_train_batch_size=bs, per_device_eval_batch_size=bs*,\n    num_train_epochs=epochs, weight_decay=, report_to=, seed=random_seed)\n\nmodel = AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=)\n</code></pre>\n<p>The argument <code>num_labels=1</code> configures the final linear layer of the model to have a single output neuron which is appropriate for a regression task, as it leads to prediction of a single continuous value.</p>\n<p>The first model was based on the <a href=\"https://huggingface.co/microsoft/deberta-v3-small\" target=\"_blank\">deberta-v3-small</a> model (<a href=\"https://openreview.net/forum?id=XPZIaotutsD\" target=\"_blank\">link to article</a>) with 44M parameters pre-trained on general text data. The \"sparse\" input used for the DeBERTa model was:</p>\n<pre><code>trdf[] =  + trdf.cell_type +  + trdf.sm_name +  + trdf.gene\n</code></pre>\n<p>Fine-tuning the DeBERTa for 20-25 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 48-60 hours (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo\" target=\"_blank\">Transformer DeBERTa (Demo) Notebook</a> for a training demo that runs on Kaggle in ~2 hours).</p>\n<p>The second model was based on the <a href=\"https://huggingface.co/nlpie/tiny-biobert\" target=\"_blank\">tiny-biobert</a> model (<a href=\"https://doi.org/10.48550/arxiv.2209.03182\" target=\"_blank\">link to article</a>) with 15M parameters pre-trained on citations and abstracts of biomedical literature using the PubMed dataset. The \"verbose\" input used for the TinyBioBert model was:</p>\n<pre><code>trdf[] =  + trdf.gene +  + trdf.cell_type +  + trdf.sm_name + \n</code></pre>\n<p>Fine-tuning the TinyBioBert model for 20-30 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 24-30 hours (see <a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo\" target=\"_blank\">Transformer TinyBioBert (Demo) Notebook</a> for a training demo that runs on Kaggle in ~1:30 hours).</p>\n<p>Interestingly, the \"sparse\" input used for the DeBERTa model did not lead to very good performance in the TinyBioBert model, whereas the more \"verbose\" input used for the TinyBioBert model led to slower training and worse results in the DeBERTa model. This is reflected in the following submission scores (all trained for 5 epochs using the same random seed of 42):</p>\n<pre><code>                        public LB     privateLB\nDeBERTa sparse:                      \nDeBERTa verbose:                     \nTinyBioBert sparse:                  \nTinyBioBert verbose:                 \n</code></pre>\n<p>The transformer part of the final ensemble consisted of four TinyBioBert models. These were trained for 30 epochs using seed 42, 30 epochs using seed 546, 20 epochs using seed 419, and 5x5 epochs using different random seeds. Each model was assigned a weight of 0.15 in the Transformer sub-ensemble. The single DeBERTa model used was trained for a total of 22 epochs using different random seeds, and was assigned a weight of 0.4.</p>\n<h2>4. Robustness</h2>\n<p>The robustness of each individual model type was evaluated using triplicate submissions with different random seeds by calculating the mean and standard deviation of the scores achieved on the public and private leaderboard (in the case of transformers, the models also varied slightly in the number of epochs trained). </p>\n<p>TabMod NN (600): 0.567 ± 0.003 public, 0.777 ± 0.007 private.<br>\nTabMod NN (1000): 0.562 ± 0.002 public, 0.778 ± 0.003 private</p>\n<p>Random Forest: 0.592 ± 0.004 public, 0.781 ± 0.002 private.</p>\n<p>XGBoosted Forest: 0.588 ± 0.002 public, 0.780 ± 0.001 private.</p>\n<p>Transformer (DeBERTa) 22-25 epochs: 0.606 ± 0.013 public, 0.784 ± 0.007 private.<br>\nTransformer (TinyBioBert) 25-30 epochs: 0.607 ± 0.003 public, 0.777 ± 0.004 private.</p>\n<p>These results show that the fine-tuned DeBERTa model is the least robust of the models. The fine-tuned TinyBioBert model consistently performed as well as the TabMod NNs on the private leaderboard, but consistently looked worse on the public leaderboard.</p>\n<h2>5. Documentation &amp; code style</h2>\n<p>The code is documented in the notebooks linked throughout the notebook. Additionally, a list of links to all notebooks and datasets can be found in Section 6.</p>\n<h2>6. Reproducibility</h2>\n<p>The source code and further documentation for all models as well as additional datasets can be accessed through the following links:</p>\n<p>Model Notebooks:<br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising\" target=\"_blank\">TabMod NN Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds\" target=\"_blank\">Random Forest Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds\" target=\"_blank\">XGBoosted Forest Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo\" target=\"_blank\">Transformer DeBERTa (Demo) Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo\" target=\"_blank\">Transformer TinyBioBert (Demo) Notebook</a></p>\n<p>Other Notebooks:<br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings\" target=\"_blank\">Look at Embeddings Notebook</a><br>\n<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions\" target=\"_blank\">Prediction Correlation Notebook</a></p>\n<p>Datasets:<br>\n<a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features\" target=\"_blank\">OP2 Additional Features Dataset</a><br>\n<a href=\"https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings\" target=\"_blank\">OP2 TabMod NN Embeddings Dataset</a></p>",
      "rawMarkdown": "Thank you to the organizers of the Open Problems – Single-Cell Perturbations competition for this incredible opportunity. \n\nMy approach combines four model types: Tabular Models with embeddings and dense layers, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests, all trained on a \"melted\" format of the training data. \n\nIn the following sections, I provide a comprehensive breakdown of my approach and the specific models employed. Links to models and other resources are provided throughout the text and can also be found in Section 6.\n\n~~Please note that my solution write-up was too long to post in a single part. Therefore, I have added the second part, covering Sections 3 to 6, in a comment below.~~\n## 1. Integration of biological knowledge\n\nMy initial approach involved training Tabular Neural Networks and Transformers specifically on the categorical features available in the training data. These models excel in encoding complex relationships within their embedding layers. Concurrently, I planned to augment the training data with information sourced from biological databases and web searches, and I intended to incorporate calculated molecular descriptors derived from SMILES. This enriched dataset would then be used to train tree-based models.\n### Molecular Descriptors\n\nI calculated 1600+ molecular descriptors using the [mordred](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-018-0258-y) python library and trained a random forest using all molecular descriptors in order to select the most important features (importance > 0.5 %) based on the feature importance results of the training run, which yielded a set of 23 molecular descriptors. All molecular descriptors as well as the list of most important descriptors can be found in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features) in the files ```mol_descriptors.parquet``` and ```important_mol_descriptors.csv```, respectively.\n```python\nimport rdkit, rdkit.Chem\nimport mordred, mordred.descriptors\n\ncalc = mordred.Calculator(mordred.descriptors, ignore_3D=True)\nmolecules = [rdkit.Chem.MolFromSmiles(smi) for smi in compounds['SMILES']]\nfeatures = calc.pandas(molecules)\n```\nHere ```compounds['SMILES']``` contained a list of unique SMILES strings from the training data.\n### Gene Information\n\nGene information was obtained from the ENSEMBL database through the [pybiomart](https://jrderuiter.github.io/pybiomart/index.html) python package. Note that gene info could only be obtained for 13,423 of the 18,211 genes in this way.\n\n```python\nfrom pybiomart import Server\n\nserver = Server(host='http://www.ensembl.org')\ndataset = server['ENSEMBL_MART_ENSEMBL']['hsapiens_gene_ensembl']\nresults = dataset.query(attributes=['ensembl_gene_id', 'external_gene_name', 'cdna', 'start_position', 'end_position'], filters={'link_ensembl_gene_id': id_filter, 'transcript_is_canonical': True})\n```\n\nFrom there I added additional features. For example, the number of each nucleobase (ACGT) obtained from the sequence data seemed particularly useful, according to feature importance results evaluated after training of Random Forests. A data frame containing added gene info can be found in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features) as ```geneinfo.parquet```.\n### Cell Features\n\nFinding useful features that distinguish the 6 different cell types in the training data was a challenge for me. I ended up browsing Wikipedia, PubMed, and the [website](https://www.immunology.org/public-information/bitesized-immunology/cells) of the British Society for Immunology for information that I could add as cell features. The resulting data frame can be found in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features) as ```cellinfo.csv```.\n### Result of Direct Integration of Biological Knowledge\n\nThe best result that I was able to obtain using a tree-based model after adding the features described above was a Random Forest that achieved scores of 0.628 and 0.836 on the public and private leaderboards, respectively. For comparison, my best Random Forest trained on embeddings learned by a Tabular Neural Network (TabMod NN) achieved scores of 0.588 and 0.780.\n### Embeddings\n\nDue to the inadequacy of direct integration of molecular descriptors, gene, and cell information, I decided to train Random and XGBoosted Forests on embeddings learned by the TabMod Neural Network. The use of embeddings significantly improved my results and might reasonably be considered an indirect form of \"integration of biological knowledge,\" because neural networks like TabMod NN are able to uncover and encode complex, often non-linear relationships inherent in biological systems into these embedding vectors. The embeddings obtained from TabMod NN can be found in the [OP2 TabMod NN Embeddings Dataset](https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings).\n\nPlots of the ```cell_type```, ```sm_name```, and ```gene``` embeddings after PCA using 2 components as well as the corresponding code can be found in the [Look at Embeddings Notebook](https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings). The plot of the ```gene``` embeddings shows two large distinct clusters – one very dense, the other less dense. The red marks with labels show a random selection of gene names to be displayed.\n\n![Gene Embeddings](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2Fa3a4ee89d9b9f0bc7d51f7c55e6faad6%2Fgene_embeds.png?generation=1702353451391406&alt=media)\n\n## 2. Exploration of the problem\n\nAt first glance, the training data seemed like a collaborative filtering problem, where cell types are the users, drugs are the items, and gene expression confidence values are the targets. However, thinking about it that way leads to 18211 collaborative filtering problems – one for every gene – and I was intimidated by the idea of training 18211 models to come up with a single submission. Then I realized that collaborative filtering is just a special case of a tabular modeling problem with two features and one target, which led to the idea to \"melt\" the training data and the test data (convert from wide to long format):\n\n```python\ntrain_df = df_de_train.melt(id_vars=['cell_type', 'sm_name'], value_vars=df_de_train.iloc[:,5:].columns, var_name='gene', value_name='value')\n```\n\nMelting of the training data results in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (```cell_type```, ```sm_name``` and ```gene```), and one target (signed -log(p-value) called ```value```). I used this data format for all my models, which then had to predict only one target for each ```cell_type```/```sm_name```/```gene``` combination. Inference using a model trained on data in that format yielded a list of 4,643,805 prediction values, which I just reshaped back into the submission format of 255 x 18,211 (see e.g. my [Random Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds)).\n## 3. Model design\n\nFor final submission, I used an ensemble of 4 different model types: Tabular Model Neural Networks, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests. These are described in more detail below. The ensemble that achieved the placement in position 30 of the private leader board (with scores of 0.553 and 0.753 in the public and private leaderboards, respectively) had the following structure:\n\n5 TabMod NNs using 600 dimensions for gene embeddings and random seeds 42, 55, 120, 457, and 736 and 5 TabMod NNs using 1000 dimensions for gene embeddings and random seeds 42, 199, 550, 855, and 970:\n\n```python\ndf1 = df1_1*0.1 + df1_2*0.1 + df1_3*0.1 + df1_4*0.1 + df1_5*0.1 + df1_6*0.1 + df1_7*0.1 + df1_8*0.1 + df1_9*0.1 + df1_10*0.1\n```\n\n5 Random Forests using random seeds 209, 569, 739, 885, and 926:\n\n```python\ndf2 = df2_1*0.2 + df2_2*0.2 + df2_3*0.2 + df2_4*0.2 + df2_5*0.2\n```\n\n5 XGBoosted Forests using random seeds 117, 150, 234, 624, and 804:\n\n```python\ndf3 = df3_1*0.2 + df3_2*0.2 + df3_3*0.2 + df3_4*0.2 + df3_5*0.2\n```\n\n5 fine-tuned Transformer Models, TinyBioBert and DeBERTa-v3-small (see details below):\n\n```python\ndf4 = df4_1*0.15 + df4_2*0.15 + df4_3*0.15 + df4_4*0.15 + df4_5*0.4\n```\n\nThe weights used for the final submission were:\n\n```python\nsubmission = df1*0.6 + df2*0.1 + df3*0.1 + df4*0.2\n```\n\nThe following figure shows a correlation heat map of the predictions of the different model types (see the [Prediction Correlation Notebook](https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions) for the source code):\n\n![Prediction Correlation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F261951234c01103ce2cb74c09c7dad41%2Fprediction_correlation.png?generation=1702353628253144&alt=media)\n\nIn the subsequent sections I present detailed descriptions of each model type, along with links to the relevant Kaggle notebooks. These notebooks are designed to run on Kaggle. However, the fine-tuning of transformer models, utilizing the full training data, was conducted on SaturnCloud. To adapt these notebooks for Kaggle, the training data was substantially reduced. As a result, the transformer models showcased on Kaggle are for demonstrative purposes only and do not achieve scores indicative of their full potential in the competition.\n### TabMod NN\n\nTabMod NN performance was best when the training data was denoised using PCA and 10 components prior to training (see the [TabMod NN Notebook](https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising) for details and code). The two kinds of tabular neural networks used for final submission were identical except for the dimensionality (600 and 1000) of gene embeddings used. The models were based on fast.ai’s ```tabular_learner``` with the following configuration:\n\n```python\nlearn = tabular_learner(dls, y_range=(y_min, y_max), emb_szs=emb_szs, layers=[1000, 500, 250], n_out=1, loss_func=F.mse_loss)\n```\n\nTo achieve a gene embedding size of 1000 instead of the default 389, the ```emb_szs``` dictionary was customized.\n\n```python\nemb_szs = {'cell_type': 5, 'sm_name': 26, 'gene': 1000}\n```\n\nThe range of the final sigmoid layer was set to ```y_min``` to ```y_max``` which were obtained by determining the minimum and maximum target values in the training data after PCA denoising. The three dense layers of size 1000, 500, and 250 were optimized for appropriate expressivity. Finally, the mean squared loss function was chosen to best reflect the evaluation in the leaderboard. Experiments using mean absolute error, huber loss, and log-cosh did not lead to improvements of the model and instead reduced model performance.\n### Random Forest\n\nRandom Forests (see [Random Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds)) were trained after adding embeddings for ```cell_type```, ```sm_name```, and ```gene``` to the training data. These embeddings were obtained from a TabMod NN similar to the one described in the previous section with gene embeddings of dimensionality 1000 but without the use of PCA denoising on the training data (see [Export Embeddings Notebook](https://www.kaggle.com/code/frenio/30-op2scp-export-tabmod-nn-embeddings) for the code). Of those the ```cell_type``` embeddings of dimensionality 5 were left unchanged, whereas ```sm_name``` and ```gene``` embeddings (of dimensionality 26 and 1000, respectively) were reduced to 10 components each using PCA.\n\nAll Random Forest models used in the final submission were identical except for the random seeds used. They were based on scikitlearn’s ```RandomForestRegressor``` using 100 trees on 66 % of samples in order to achieve a meaningful out-of-bag error and score. The square root of feature number was used as the maximum number of features per tree which has been [shown](https://scikit-learn.org/stable/auto_examples/ensemble/plot_ensemble_oob.html) to be beneficial for the model’s ability to generalize. The minimum number of samples per leaf was set to 5 in order to achieve high expressivity. Such a Random Forest could be trained using the following example code:\n\n```python\nfrom sklearn.ensemble import RandomForestRegressor\n\nm = RandomForestRegressor(n_jobs=-1, n_estimators=100, max_samples=0.66, max_features='sqrt', min_samples_leaf=5, oob_score=True)\nm.fit(xs, y)\n```\n### XGBoosted Forest\n\nThe same embeddings used to train Random Forests were also used to train XGBoosted Forests using the same dimensionality reductions.\n\nThe XGBoosted Forests used in the final submission were also identical to each other except for the random seed used. They were trained using the following code (see [XGBoosted Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds) for details):\n\n```python\nimport xgboost as xgb\n\nm = xgb.XGBRegressor(device='cuda', n_estimators=185, learning_rate=0.01, max_depth=20, min_child_weight=1, gamma=0, subsample=1, colsample_bytree=0.55)\nm.fit(xs, y)\n```\n\nThe model parameters were optimized using the cross validation scheme proposed by [AmbrosM](https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&cellId=8), where all but 10 % of the data of one of four cell types is used as a validation set. It should be noted that the performance of the model could potentially be enhanced by configuring the ```max_depth``` parameter to values exceeding 20. This approach was not extensively explored due to the higher computational costs associated with larger values.\n### Transformer Models\n\nTwo kinds of Transformer models from the [Huggingface Transformers](https://huggingface.co/docs/transformers/index) data base were fine-tuned to output a number when given a standardized input sentence. The general model configuration is shown in the following code snippet:\n\n```python\nfrom transformers import TrainingArguments, Trainer\n\nargs = TrainingArguments('outputs', save_steps=steps, learning_rate=8e-5, warmup_ratio=0.1, lr_scheduler_type='cosine', fp16=True,\n    evaluation_strategy=\"epoch\", per_device_train_batch_size=bs, per_device_eval_batch_size=bs*2,\n    num_train_epochs=epochs, weight_decay=0.01, report_to='none', seed=random_seed)\n\nmodel = AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=1)\n```\n\nThe argument ```num_labels=1``` configures the final linear layer of the model to have a single output neuron which is appropriate for a regression task, as it leads to prediction of a single continuous value.\n\nThe first model was based on the [deberta-v3-small](https://huggingface.co/microsoft/deberta-v3-small) model ([link to article](https://openreview.net/forum?id=XPZIaotutsD)) with 44M parameters pre-trained on general text data. The \"sparse\" input used for the DeBERTa model was:\n```python\ntrdf['input'] = 'CELL TYPE: ' + trdf.cell_type + '; TREATMENT: ' + trdf.sm_name + '; GENE: ' + trdf.gene\n```\nFine-tuning the DeBERTa for 20-25 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 48-60 hours (see [Transformer DeBERTa (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo) for a training demo that runs on Kaggle in ~2 hours).\n\nThe second model was based on the [tiny-biobert](https://huggingface.co/nlpie/tiny-biobert) model ([link to article](https://doi.org/10.48550/arxiv.2209.03182)) with 15M parameters pre-trained on citations and abstracts of biomedical literature using the PubMed dataset. The \"verbose\" input used for the TinyBioBert model was:\n```python\ntrdf['input'] = \"Estimate the −log(p-value) confidence for change in gene expression of \" + trdf.gene + \" in \" + trdf.cell_type + \" when treated with \" + trdf.sm_name + \" compared to DMSO.\"\n```\nFine-tuning the TinyBioBert model for 20-30 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 24-30 hours (see [Transformer TinyBioBert (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo) for a training demo that runs on Kaggle in ~1:30 hours).\n\nInterestingly, the \"sparse\" input used for the DeBERTa model did not lead to very good performance in the TinyBioBert model, whereas the more \"verbose\" input used for the TinyBioBert model led to slower training and worse results in the DeBERTa model. This is reflected in the following submission scores (all trained for 5 epochs using the same random seed of 42):\n\n```python\n                        public LB     privateLB\nDeBERTa sparse:             0.630         0.807\nDeBERTa verbose:            0.637         0.810\nTinyBioBert sparse:         0.634         0.812\nTinyBioBert verbose:        0.624         0.807\n```\n\nThe transformer part of the final ensemble consisted of four TinyBioBert models. These were trained for 30 epochs using seed 42, 30 epochs using seed 546, 20 epochs using seed 419, and 5x5 epochs using different random seeds. Each model was assigned a weight of 0.15 in the Transformer sub-ensemble. The single DeBERTa model used was trained for a total of 22 epochs using different random seeds, and was assigned a weight of 0.4.\n## 4. Robustness\n\nThe robustness of each individual model type was evaluated using triplicate submissions with different random seeds by calculating the mean and standard deviation of the scores achieved on the public and private leaderboard (in the case of transformers, the models also varied slightly in the number of epochs trained). \n\nTabMod NN (600): 0.567 ± 0.003 public, 0.777 ± 0.007 private.\nTabMod NN (1000): 0.562 ± 0.002 public, 0.778 ± 0.003 private\n\nRandom Forest: 0.592 ± 0.004 public, 0.781 ± 0.002 private.\n\nXGBoosted Forest: 0.588 ± 0.002 public, 0.780 ± 0.001 private.\n\nTransformer (DeBERTa) 22-25 epochs: 0.606 ± 0.013 public, 0.784 ± 0.007 private.\nTransformer (TinyBioBert) 25-30 epochs: 0.607 ± 0.003 public, 0.777 ± 0.004 private.\n\nThese results show that the fine-tuned DeBERTa model is the least robust of the models. The fine-tuned TinyBioBert model consistently performed as well as the TabMod NNs on the private leaderboard, but consistently looked worse on the public leaderboard.\n## 5. Documentation & code style\n\nThe code is documented in the notebooks linked throughout the notebook. Additionally, a list of links to all notebooks and datasets can be found in Section 6.\n## 6. Reproducibility\n\nThe source code and further documentation for all models as well as additional datasets can be accessed through the following links:\n\nModel Notebooks:\n[TabMod NN Notebook](https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising)\n[Random Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds)\n[XGBoosted Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds)\n[Transformer DeBERTa (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo)\n[Transformer TinyBioBert (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo)\n\nOther Notebooks:\n[Look at Embeddings Notebook](https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings)\n[Prediction Correlation Notebook](https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions)\n\nDatasets:\n[OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features)\n[OP2 TabMod NN Embeddings Dataset](https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings)",
      "votes": 8
    },
    {
      "id": 2562757,
      "postDate": "2023-12-15T16:33:55.767Z",
      "content": "<p>This comment used to contain Sections 3 to 6 of my solution write-up because a bug kept me from posting it in whole. Now that the bug is fixed, I replaced the original comment with this notice and added Sections 3 to 6 to the main topic.</p>",
      "rawMarkdown": "This comment used to contain Sections 3 to 6 of my solution write-up because a bug kept me from posting it in whole. Now that the bug is fixed, I replaced the original comment with this notice and added Sections 3 to 6 to the main topic.",
      "votes": 5
    },
    {
      "id": 2569948,
      "postDate": "2023-12-21T17:36:16.460Z",
      "content": "<p>Thanks for sharing, Frenio! Your score of 0.553 on the leaderboard really gave us a lot of nightmares. You led this challenge for a very long time!</p>",
      "rawMarkdown": "Thanks for sharing, Frenio! Your score of 0.553 on the leaderboard really gave us a lot of nightmares. You led this challenge for a very long time!",
      "votes": 3,
      "replies": [
        {
          "id": 2570004,
          "postDate": "2023-12-21T19:21:04.933Z",
          "content": "<p>Haha, until you came in with your 0.546 score and gave me nightmares! And towards the end, before overfitting public notebooks took over, <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> caused nightmares for us all I suppose. </p>",
          "rawMarkdown": "Haha, until you came in with your 0.546 score and gave me nightmares! And towards the end, before overfitting public notebooks took over, @eliork caused nightmares for us all I suppose. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 2562953,
      "postDate": "2023-12-15T20:20:42.990Z",
      "content": "<p>Congratulations and thanks for sharing your precious insights (topic and comments above)</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your precious insights (topic and comments above)",
      "votes": 3
    },
    {
      "id": 2562826,
      "postDate": "2023-12-15T17:18:23.063Z",
      "content": "<p>Thanks a lot for sharing ! <br>\nAny ideas what can be interpretation of the two clusters for genes ? <br>\nTo my mind there is no two natural clusters in human genes. There can be be different approaches giving different results , several are in HPA like: <a href=\"https://www.proteinatlas.org/humanproteome/tissue/expression+cluster\" target=\"_blank\">https://www.proteinatlas.org/humanproteome/tissue/expression+cluster</a><br>\nPS<br>\nWould you be so kind to make your key submit open , if possible ? </p>",
      "rawMarkdown": "Thanks a lot for sharing ! \nAny ideas what can be interpretation of the two clusters for genes ? \nTo my mind there is no two natural clusters in human genes. There can be be different approaches giving different results , several are in HPA like: https://www.proteinatlas.org/humanproteome/tissue/expression+cluster\nPS\nWould you be so kind to make your key submit open , if possible ? \n",
      "votes": 4,
      "replies": [
        {
          "id": 2562946,
          "postDate": "2023-12-15T20:10:47.260Z",
          "content": "<p>Thank you for reading and for providing so many useful public notebooks to get people started in this competition! </p>\n<p>I do not know how to interpret these two gene clusters, unfortunately, but do I recall correctly that you also found two clusters of genes in one of your earlier public EDA notebooks using a different clustering approach? (I wasn't able to find it just now.) Happy to discuss more if you have ideas.</p>\n<p>PS: I am not sure I understand what you mean by key submit, but if your interested in my final submission notebook, I just made it public and you can find it <a href=\"https://www.kaggle.com/code/frenio/open-problems-super-ensemble-submission/notebook\" target=\"_blank\">here</a>. I only used that notebook to load and manually blend the individual model predictions, though. Please let me know if I can provide any other resources that might be helpful.</p>",
          "rawMarkdown": "Thank you for reading and for providing so many useful public notebooks to get people started in this competition! \n\nI do not know how to interpret these two gene clusters, unfortunately, but do I recall correctly that you also found two clusters of genes in one of your earlier public EDA notebooks using a different clustering approach? (I wasn't able to find it just now.) Happy to discuss more if you have ideas.\n\nPS: I am not sure I understand what you mean by key submit, but if your interested in my final submission notebook, I just made it public and you can find it [here](https://www.kaggle.com/code/frenio/open-problems-super-ensemble-submission/notebook). I only used that notebook to load and manually blend the individual model predictions, though. Please let me know if I can provide any other resources that might be helpful.",
          "votes": 3,
          "replies": [
            {
              "id": 2562959,
              "postDate": "2023-12-15T20:30:50.890Z",
              "content": "<p>Thanks for kind words ! </p>\n<p>Indeed in here there are something like 2 clusters in genes (one much bigger than the other) here:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=143061639&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=143061639&amp;cellId=21</a><br>\n(it works quite long time, so in other version of the notebook I reduced number of genes and we do not see that figure, so it is only in version 19).<br>\nI did not analyzed that much… Thanks for reminding me that..</p>",
              "rawMarkdown": "Thanks for kind words ! \n\nIndeed in here there are something like 2 clusters in genes (one much bigger than the other) here:\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=143061639&cellId=21\n(it works quite long time, so in other version of the notebook I reduced number of genes and we do not see that figure, so it is only in version 19).\nI did not analyzed that much... Thanks for reminding me that..\n",
              "votes": 2
            },
            {
              "id": 2562961,
              "postDate": "2023-12-15T20:35:40.290Z",
              "content": "<p>Great, thank you for finding the version that I had in mind!</p>",
              "rawMarkdown": "Great, thank you for finding the version that I had in mind!",
              "votes": 2
            }
          ]
        },
        {
          "id": 2564213,
          "postDate": "2023-12-16T22:41:56.767Z",
          "content": "<p>There is interesting picture from <a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459623\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459623</a><br>\nHis interpretation - up vs down regulated genes.  Seems very reasonable. May be it is similar to yours. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6f8e437e2b8421895620db4fcdb68914%2Fphoto_2023-12-16_23-39-01.jpg?generation=1702766361769740&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "There is interesting picture from @zhijianli https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459623\nHis interpretation - up vs down regulated genes.  Seems very reasonable. May be it is similar to yours. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6f8e437e2b8421895620db4fcdb68914%2Fphoto_2023-12-16_23-39-01.jpg?generation=1702766361769740&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2564263,
              "postDate": "2023-12-17T01:16:23.747Z",
              "content": "<p>Thank you for pointing me to this! I'll look into his solution more closely.</p>",
              "rawMarkdown": "Thank you for pointing me to this! I'll look into his solution more closely.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2728295,
      "postDate": "2024-04-02T07:03:21.113Z",
      "content": "<p>Hi, interesting and well described solution with melting the data and learning the embedding representations, thanks for sharing! In the Random Forest model, you used embeddings from a TabMod NN with a dimensionality of 1000 for genes and without using denoising on the training data, and as I understand it, it was better, while the TabMod NN itself performed better with embeddings after denoising and with a dimensionality of 600.</p>\n<p>Maybe it's just my poor understanding, but as I understand it, the better the model performs, the better the quality of the embeddings it learns during training. So how would you explain that lower quality embeddings performed better in RF?</p>",
      "rawMarkdown": "Hi, interesting and well described solution with melting the data and learning the embedding representations, thanks for sharing! In the Random Forest model, you used embeddings from a TabMod NN with a dimensionality of 1000 for genes and without using denoising on the training data, and as I understand it, it was better, while the TabMod NN itself performed better with embeddings after denoising and with a dimensionality of 600.\n\nMaybe it's just my poor understanding, but as I understand it, the better the model performs, the better the quality of the embeddings it learns during training. So how would you explain that lower quality embeddings performed better in RF?",
      "votes": 2,
      "replies": [
        {
          "id": 2754062,
          "postDate": "2024-04-15T19:17:38.460Z",
          "content": "<p>Thank you for the feedback! I was definitely surprised when I found that the RF performs better with the embeddings learned from the plain data without denoising, but unfortunately haven't come up with a good explanation so far. I would be very interested to hear if you have any hypotheses. On the subject of embedding sizes, I think the differences in performance between the dimensionalities of 600 and 1000 weren't significant.</p>",
          "rawMarkdown": "Thank you for the feedback! I was definitely surprised when I found that the RF performs better with the embeddings learned from the plain data without denoising, but unfortunately haven't come up with a good explanation so far. I would be very interested to hear if you have any hypotheses. On the subject of embedding sizes, I think the differences in performance between the dimensionalities of 600 and 1000 weren't significant.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2579010,
      "postDate": "2023-12-29T17:36:30.647Z",
      "content": "<p>Thanks again for the great write-up !</p>\n<p>I wonder about the \"Gene Information\" section - you mention that you convert genes information into features. <br>\nWould you be so kind to clarify it a bit - because genes -are our targets  - so any info on genes -  do not depend on our features: cell type/drug, so it is a kind constant across sample-wise direction. There are enormous amounts   of information on human genes, but typically there is no way to use it in such kind of tasks, because it is \"constant\" across cells/cell types/drugs direction, so not a feature.</p>",
      "rawMarkdown": "Thanks again for the great write-up !\n\nI wonder about the \"Gene Information\" section - you mention that you convert genes information into features. \nWould you be so kind to clarify it a bit - because genes -are our targets  - so any info on genes -  do not depend on our features: cell type/drug, so it is a kind constant across sample-wise direction. There are enormous amounts   of information on human genes, but typically there is no way to use it in such kind of tasks, because it is \"constant\" across cells/cell types/drugs direction, so not a feature.\n",
      "votes": 2,
      "replies": [
        {
          "id": 2579263,
          "postDate": "2023-12-30T00:28:09.240Z",
          "content": "<p>Thank you for keeping up the research on this problem! I am more than happy to try and clarify. </p>\n<p>As I wrote in \"Exploration of the Problem\" I converted all input data into long format such that genes actually turned into a feature. Here is a screenshot of the \"melted\" data frame to make it more concrete:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F52cac9c1ec9653eb971174281a0cef73%2FMeltedTrainData.png?generation=1703893251100013&amp;alt=media\"></p>\n<p>You can find the additional gene features I extracted from the ENSEMBL database in the <code>geneinfo.parquet</code> file in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features?select=geneinfo.parquet\" target=\"_blank\">OP2 Additional Features Dataset</a> that I've made public. I merged this geneinfo with the training data on <code>gene</code> like shown in the following code snippet:</p>\n<pre><code>train_df_ginf = pd.merge(train_df, selected_geneinfo, on=, how=)\n</code></pre>\n<p>Please note that the CPU RAM on Kaggle is too low for this step to work, so this needs to be done on a different platform. I just published a notebook that shows how I added all additional features to the data (<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-randomforest-with-additional-features/notebook\" target=\"_blank\">Random Forest with Additional Features</a>), but due to limited RAM it will not run on Kaggle, unfortunately.</p>\n<p>I hope this helps and I am happy to explain more, so let me know if you have further questions.</p>",
          "rawMarkdown": "Thank you for keeping up the research on this problem! I am more than happy to try and clarify. \n\nAs I wrote in \"Exploration of the Problem\" I converted all input data into long format such that genes actually turned into a feature. Here is a screenshot of the \"melted\" data frame to make it more concrete:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F52cac9c1ec9653eb971174281a0cef73%2FMeltedTrainData.png?generation=1703893251100013&alt=media\" width=\"329\" height=\"300\">\n\nYou can find the additional gene features I extracted from the ENSEMBL database in the ```geneinfo.parquet``` file in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features?select=geneinfo.parquet) that I've made public. I merged this geneinfo with the training data on ```gene``` like shown in the following code snippet:\n\n```python\ntrain_df_ginf = pd.merge(train_df, selected_geneinfo, on='gene', how='left')\n```\n\nPlease note that the CPU RAM on Kaggle is too low for this step to work, so this needs to be done on a different platform. I just published a notebook that shows how I added all additional features to the data ([Random Forest with Additional Features](https://www.kaggle.com/code/frenio/30-op2scp-randomforest-with-additional-features/notebook)), but due to limited RAM it will not run on Kaggle, unfortunately.\n\nI hope this helps and I am happy to explain more, so let me know if you have further questions.\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 2562754,
      "postDate": "2023-12-15T16:31:33.333Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2562757,
      "author_name": "Frenio Redeker",
      "author_url": "",
      "post_date": "2023-12-15T16:33:55.767000",
      "content": "<p>This comment used to contain Sections 3 to 6 of my solution write-up because a bug kept me from posting it in whole. Now that the bug is fixed, I replaced the original comment with this notice and added Sections 3 to 6 to the main topic.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2569948,
      "author_name": "Pablo Rodriguez-Mier",
      "author_url": "",
      "post_date": "2023-12-21T17:36:16.460000",
      "content": "<p>Thanks for sharing, Frenio! Your score of 0.553 on the leaderboard really gave us a lot of nightmares. You led this challenge for a very long time!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2570004,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2023-12-21T19:21:04.933000",
          "content": "<p>Haha, until you came in with your 0.546 score and gave me nightmares! And towards the end, before overfitting public notebooks took over, <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> caused nightmares for us all I suppose. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2562953,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2023-12-15T20:20:42.990000",
      "content": "<p>Congratulations and thanks for sharing your precious insights (topic and comments above)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2562826,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2023-12-15T17:18:23.063000",
      "content": "<p>Thanks a lot for sharing ! <br>\nAny ideas what can be interpretation of the two clusters for genes ? <br>\nTo my mind there is no two natural clusters in human genes. There can be be different approaches giving different results , several are in HPA like: <a href=\"https://www.proteinatlas.org/humanproteome/tissue/expression+cluster\" target=\"_blank\">https://www.proteinatlas.org/humanproteome/tissue/expression+cluster</a><br>\nPS<br>\nWould you be so kind to make your key submit open , if possible ? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2562946,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2023-12-15T20:10:47.260000",
          "content": "<p>Thank you for reading and for providing so many useful public notebooks to get people started in this competition! </p>\n<p>I do not know how to interpret these two gene clusters, unfortunately, but do I recall correctly that you also found two clusters of genes in one of your earlier public EDA notebooks using a different clustering approach? (I wasn't able to find it just now.) Happy to discuss more if you have ideas.</p>\n<p>PS: I am not sure I understand what you mean by key submit, but if your interested in my final submission notebook, I just made it public and you can find it <a href=\"https://www.kaggle.com/code/frenio/open-problems-super-ensemble-submission/notebook\" target=\"_blank\">here</a>. I only used that notebook to load and manually blend the individual model predictions, though. Please let me know if I can provide any other resources that might be helpful.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2562959,
              "author_name": "Alexander Chervov",
              "author_url": "",
              "post_date": "2023-12-15T20:30:50.890000",
              "content": "<p>Thanks for kind words ! </p>\n<p>Indeed in here there are something like 2 clusters in genes (one much bigger than the other) here:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=143061639&amp;cellId=21\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s?scriptVersionId=143061639&amp;cellId=21</a><br>\n(it works quite long time, so in other version of the notebook I reduced number of genes and we do not see that figure, so it is only in version 19).<br>\nI did not analyzed that much… Thanks for reminding me that..</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2562961,
              "author_name": "Frenio Redeker",
              "author_url": "",
              "post_date": "2023-12-15T20:35:40.290000",
              "content": "<p>Great, thank you for finding the version that I had in mind!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2564213,
          "author_name": "Alexander Chervov",
          "author_url": "",
          "post_date": "2023-12-16T22:41:56.767000",
          "content": "<p>There is interesting picture from <a href=\"https://www.kaggle.com/zhijianli\" target=\"_blank\">@zhijianli</a> <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459623\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/459623</a><br>\nHis interpretation - up vs down regulated genes.  Seems very reasonable. May be it is similar to yours. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262596%2F6f8e437e2b8421895620db4fcdb68914%2Fphoto_2023-12-16_23-39-01.jpg?generation=1702766361769740&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2564263,
              "author_name": "Frenio Redeker",
              "author_url": "",
              "post_date": "2023-12-17T01:16:23.747000",
              "content": "<p>Thank you for pointing me to this! I'll look into his solution more closely.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2728295,
      "author_name": "Antonina Dolgorukova",
      "author_url": "",
      "post_date": "2024-04-02T07:03:21.113000",
      "content": "<p>Hi, interesting and well described solution with melting the data and learning the embedding representations, thanks for sharing! In the Random Forest model, you used embeddings from a TabMod NN with a dimensionality of 1000 for genes and without using denoising on the training data, and as I understand it, it was better, while the TabMod NN itself performed better with embeddings after denoising and with a dimensionality of 600.</p>\n<p>Maybe it's just my poor understanding, but as I understand it, the better the model performs, the better the quality of the embeddings it learns during training. So how would you explain that lower quality embeddings performed better in RF?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2754062,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2024-04-15T19:17:38.460000",
          "content": "<p>Thank you for the feedback! I was definitely surprised when I found that the RF performs better with the embeddings learned from the plain data without denoising, but unfortunately haven't come up with a good explanation so far. I would be very interested to hear if you have any hypotheses. On the subject of embedding sizes, I think the differences in performance between the dimensionalities of 600 and 1000 weren't significant.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2579010,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2023-12-29T17:36:30.647000",
      "content": "<p>Thanks again for the great write-up !</p>\n<p>I wonder about the \"Gene Information\" section - you mention that you convert genes information into features. <br>\nWould you be so kind to clarify it a bit - because genes -are our targets  - so any info on genes -  do not depend on our features: cell type/drug, so it is a kind constant across sample-wise direction. There are enormous amounts   of information on human genes, but typically there is no way to use it in such kind of tasks, because it is \"constant\" across cells/cell types/drugs direction, so not a feature.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2579263,
          "author_name": "Frenio Redeker",
          "author_url": "",
          "post_date": "2023-12-30T00:28:09.240000",
          "content": "<p>Thank you for keeping up the research on this problem! I am more than happy to try and clarify. </p>\n<p>As I wrote in \"Exploration of the Problem\" I converted all input data into long format such that genes actually turned into a feature. Here is a screenshot of the \"melted\" data frame to make it more concrete:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F52cac9c1ec9653eb971174281a0cef73%2FMeltedTrainData.png?generation=1703893251100013&amp;alt=media\"></p>\n<p>You can find the additional gene features I extracted from the ENSEMBL database in the <code>geneinfo.parquet</code> file in the <a href=\"https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features?select=geneinfo.parquet\" target=\"_blank\">OP2 Additional Features Dataset</a> that I've made public. I merged this geneinfo with the training data on <code>gene</code> like shown in the following code snippet:</p>\n<pre><code>train_df_ginf = pd.merge(train_df, selected_geneinfo, on=, how=)\n</code></pre>\n<p>Please note that the CPU RAM on Kaggle is too low for this step to work, so this needs to be done on a different platform. I just published a notebook that shows how I added all additional features to the data (<a href=\"https://www.kaggle.com/code/frenio/30-op2scp-randomforest-with-additional-features/notebook\" target=\"_blank\">Random Forest with Additional Features</a>), but due to limited RAM it will not run on Kaggle, unfortunately.</p>\n<p>I hope this helps and I am happy to explain more, so let me know if you have further questions.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2562754,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-15T16:31:33.333000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2562753": "Thank you to the organizers of the Open Problems – Single-Cell Perturbations competition for this incredible opportunity. \n\nMy approach combines four model types: Tabular Models with embeddings and dense layers, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests, all trained on a \"melted\" format of the training data. \n\nIn the following sections, I provide a comprehensive breakdown of my approach and the specific models employed. Links to models and other resources are provided throughout the text and can also be found in Section 6.\n\n~~Please note that my solution write-up was too long to post in a single part. Therefore, I have added the second part, covering Sections 3 to 6, in a comment below.~~\n## 1. Integration of biological knowledge\n\nMy initial approach involved training Tabular Neural Networks and Transformers specifically on the categorical features available in the training data. These models excel in encoding complex relationships within their embedding layers. Concurrently, I planned to augment the training data with information sourced from biological databases and web searches, and I intended to incorporate calculated molecular descriptors derived from SMILES. This enriched dataset would then be used to train tree-based models.\n### Molecular Descriptors\n\nI calculated 1600+ molecular descriptors using the [mordred](https://jcheminf.biomedcentral.com/articles/10.1186/s13321-018-0258-y) python library and trained a random forest using all molecular descriptors in order to select the most important features (importance > 0.5 %) based on the feature importance results of the training run, which yielded a set of 23 molecular descriptors. All molecular descriptors as well as the list of most important descriptors can be found in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features) in the files ```mol_descriptors.parquet``` and ```important_mol_descriptors.csv```, respectively.\n```python\nimport rdkit, rdkit.Chem\nimport mordred, mordred.descriptors\n\ncalc = mordred.Calculator(mordred.descriptors, ignore_3D=True)\nmolecules = [rdkit.Chem.MolFromSmiles(smi) for smi in compounds['SMILES']]\nfeatures = calc.pandas(molecules)\n```\nHere ```compounds['SMILES']``` contained a list of unique SMILES strings from the training data.\n### Gene Information\n\nGene information was obtained from the ENSEMBL database through the [pybiomart](https://jrderuiter.github.io/pybiomart/index.html) python package. Note that gene info could only be obtained for 13,423 of the 18,211 genes in this way.\n\n```python\nfrom pybiomart import Server\n\nserver = Server(host='http://www.ensembl.org')\ndataset = server['ENSEMBL_MART_ENSEMBL']['hsapiens_gene_ensembl']\nresults = dataset.query(attributes=['ensembl_gene_id', 'external_gene_name', 'cdna', 'start_position', 'end_position'], filters={'link_ensembl_gene_id': id_filter, 'transcript_is_canonical': True})\n```\n\nFrom there I added additional features. For example, the number of each nucleobase (ACGT) obtained from the sequence data seemed particularly useful, according to feature importance results evaluated after training of Random Forests. A data frame containing added gene info can be found in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features) as ```geneinfo.parquet```.\n### Cell Features\n\nFinding useful features that distinguish the 6 different cell types in the training data was a challenge for me. I ended up browsing Wikipedia, PubMed, and the [website](https://www.immunology.org/public-information/bitesized-immunology/cells) of the British Society for Immunology for information that I could add as cell features. The resulting data frame can be found in the [OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features) as ```cellinfo.csv```.\n### Result of Direct Integration of Biological Knowledge\n\nThe best result that I was able to obtain using a tree-based model after adding the features described above was a Random Forest that achieved scores of 0.628 and 0.836 on the public and private leaderboards, respectively. For comparison, my best Random Forest trained on embeddings learned by a Tabular Neural Network (TabMod NN) achieved scores of 0.588 and 0.780.\n### Embeddings\n\nDue to the inadequacy of direct integration of molecular descriptors, gene, and cell information, I decided to train Random and XGBoosted Forests on embeddings learned by the TabMod Neural Network. The use of embeddings significantly improved my results and might reasonably be considered an indirect form of \"integration of biological knowledge,\" because neural networks like TabMod NN are able to uncover and encode complex, often non-linear relationships inherent in biological systems into these embedding vectors. The embeddings obtained from TabMod NN can be found in the [OP2 TabMod NN Embeddings Dataset](https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings).\n\nPlots of the ```cell_type```, ```sm_name```, and ```gene``` embeddings after PCA using 2 components as well as the corresponding code can be found in the [Look at Embeddings Notebook](https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings). The plot of the ```gene``` embeddings shows two large distinct clusters – one very dense, the other less dense. The red marks with labels show a random selection of gene names to be displayed.\n\n![Gene Embeddings](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2Fa3a4ee89d9b9f0bc7d51f7c55e6faad6%2Fgene_embeds.png?generation=1702353451391406&alt=media)\n\n## 2. Exploration of the problem\n\nAt first glance, the training data seemed like a collaborative filtering problem, where cell types are the users, drugs are the items, and gene expression confidence values are the targets. However, thinking about it that way leads to 18211 collaborative filtering problems – one for every gene – and I was intimidated by the idea of training 18211 models to come up with a single submission. Then I realized that collaborative filtering is just a special case of a tabular modeling problem with two features and one target, which led to the idea to \"melt\" the training data and the test data (convert from wide to long format):\n\n```python\ntrain_df = df_de_train.melt(id_vars=['cell_type', 'sm_name'], value_vars=df_de_train.iloc[:,5:].columns, var_name='gene', value_name='value')\n```\n\nMelting of the training data results in a data frame with 11,181,554 rows (11,181,554/614 = 18,211), three features (```cell_type```, ```sm_name``` and ```gene```), and one target (signed -log(p-value) called ```value```). I used this data format for all my models, which then had to predict only one target for each ```cell_type```/```sm_name```/```gene``` combination. Inference using a model trained on data in that format yielded a list of 4,643,805 prediction values, which I just reshaped back into the submission format of 255 x 18,211 (see e.g. my [Random Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds)).\n## 3. Model design\n\nFor final submission, I used an ensemble of 4 different model types: Tabular Model Neural Networks, fine-tuned Transformer Models, Random Forests, and XGBoosted Forests. These are described in more detail below. The ensemble that achieved the placement in position 30 of the private leader board (with scores of 0.553 and 0.753 in the public and private leaderboards, respectively) had the following structure:\n\n5 TabMod NNs using 600 dimensions for gene embeddings and random seeds 42, 55, 120, 457, and 736 and 5 TabMod NNs using 1000 dimensions for gene embeddings and random seeds 42, 199, 550, 855, and 970:\n\n```python\ndf1 = df1_1*0.1 + df1_2*0.1 + df1_3*0.1 + df1_4*0.1 + df1_5*0.1 + df1_6*0.1 + df1_7*0.1 + df1_8*0.1 + df1_9*0.1 + df1_10*0.1\n```\n\n5 Random Forests using random seeds 209, 569, 739, 885, and 926:\n\n```python\ndf2 = df2_1*0.2 + df2_2*0.2 + df2_3*0.2 + df2_4*0.2 + df2_5*0.2\n```\n\n5 XGBoosted Forests using random seeds 117, 150, 234, 624, and 804:\n\n```python\ndf3 = df3_1*0.2 + df3_2*0.2 + df3_3*0.2 + df3_4*0.2 + df3_5*0.2\n```\n\n5 fine-tuned Transformer Models, TinyBioBert and DeBERTa-v3-small (see details below):\n\n```python\ndf4 = df4_1*0.15 + df4_2*0.15 + df4_3*0.15 + df4_4*0.15 + df4_5*0.4\n```\n\nThe weights used for the final submission were:\n\n```python\nsubmission = df1*0.6 + df2*0.1 + df3*0.1 + df4*0.2\n```\n\nThe following figure shows a correlation heat map of the predictions of the different model types (see the [Prediction Correlation Notebook](https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions) for the source code):\n\n![Prediction Correlation](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15827820%2F261951234c01103ce2cb74c09c7dad41%2Fprediction_correlation.png?generation=1702353628253144&alt=media)\n\nIn the subsequent sections I present detailed descriptions of each model type, along with links to the relevant Kaggle notebooks. These notebooks are designed to run on Kaggle. However, the fine-tuning of transformer models, utilizing the full training data, was conducted on SaturnCloud. To adapt these notebooks for Kaggle, the training data was substantially reduced. As a result, the transformer models showcased on Kaggle are for demonstrative purposes only and do not achieve scores indicative of their full potential in the competition.\n### TabMod NN\n\nTabMod NN performance was best when the training data was denoised using PCA and 10 components prior to training (see the [TabMod NN Notebook](https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising) for details and code). The two kinds of tabular neural networks used for final submission were identical except for the dimensionality (600 and 1000) of gene embeddings used. The models were based on fast.ai’s ```tabular_learner``` with the following configuration:\n\n```python\nlearn = tabular_learner(dls, y_range=(y_min, y_max), emb_szs=emb_szs, layers=[1000, 500, 250], n_out=1, loss_func=F.mse_loss)\n```\n\nTo achieve a gene embedding size of 1000 instead of the default 389, the ```emb_szs``` dictionary was customized.\n\n```python\nemb_szs = {'cell_type': 5, 'sm_name': 26, 'gene': 1000}\n```\n\nThe range of the final sigmoid layer was set to ```y_min``` to ```y_max``` which were obtained by determining the minimum and maximum target values in the training data after PCA denoising. The three dense layers of size 1000, 500, and 250 were optimized for appropriate expressivity. Finally, the mean squared loss function was chosen to best reflect the evaluation in the leaderboard. Experiments using mean absolute error, huber loss, and log-cosh did not lead to improvements of the model and instead reduced model performance.\n### Random Forest\n\nRandom Forests (see [Random Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds)) were trained after adding embeddings for ```cell_type```, ```sm_name```, and ```gene``` to the training data. These embeddings were obtained from a TabMod NN similar to the one described in the previous section with gene embeddings of dimensionality 1000 but without the use of PCA denoising on the training data (see [Export Embeddings Notebook](https://www.kaggle.com/code/frenio/30-op2scp-export-tabmod-nn-embeddings) for the code). Of those the ```cell_type``` embeddings of dimensionality 5 were left unchanged, whereas ```sm_name``` and ```gene``` embeddings (of dimensionality 26 and 1000, respectively) were reduced to 10 components each using PCA.\n\nAll Random Forest models used in the final submission were identical except for the random seeds used. They were based on scikitlearn’s ```RandomForestRegressor``` using 100 trees on 66 % of samples in order to achieve a meaningful out-of-bag error and score. The square root of feature number was used as the maximum number of features per tree which has been [shown](https://scikit-learn.org/stable/auto_examples/ensemble/plot_ensemble_oob.html) to be beneficial for the model’s ability to generalize. The minimum number of samples per leaf was set to 5 in order to achieve high expressivity. Such a Random Forest could be trained using the following example code:\n\n```python\nfrom sklearn.ensemble import RandomForestRegressor\n\nm = RandomForestRegressor(n_jobs=-1, n_estimators=100, max_samples=0.66, max_features='sqrt', min_samples_leaf=5, oob_score=True)\nm.fit(xs, y)\n```\n### XGBoosted Forest\n\nThe same embeddings used to train Random Forests were also used to train XGBoosted Forests using the same dimensionality reductions.\n\nThe XGBoosted Forests used in the final submission were also identical to each other except for the random seed used. They were trained using the following code (see [XGBoosted Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds) for details):\n\n```python\nimport xgboost as xgb\n\nm = xgb.XGBRegressor(device='cuda', n_estimators=185, learning_rate=0.01, max_depth=20, min_child_weight=1, gamma=0, subsample=1, colsample_bytree=0.55)\nm.fit(xs, y)\n```\n\nThe model parameters were optimized using the cross validation scheme proposed by [AmbrosM](https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&cellId=8), where all but 10 % of the data of one of four cell types is used as a validation set. It should be noted that the performance of the model could potentially be enhanced by configuring the ```max_depth``` parameter to values exceeding 20. This approach was not extensively explored due to the higher computational costs associated with larger values.\n### Transformer Models\n\nTwo kinds of Transformer models from the [Huggingface Transformers](https://huggingface.co/docs/transformers/index) data base were fine-tuned to output a number when given a standardized input sentence. The general model configuration is shown in the following code snippet:\n\n```python\nfrom transformers import TrainingArguments, Trainer\n\nargs = TrainingArguments('outputs', save_steps=steps, learning_rate=8e-5, warmup_ratio=0.1, lr_scheduler_type='cosine', fp16=True,\n    evaluation_strategy=\"epoch\", per_device_train_batch_size=bs, per_device_eval_batch_size=bs*2,\n    num_train_epochs=epochs, weight_decay=0.01, report_to='none', seed=random_seed)\n\nmodel = AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=1)\n```\n\nThe argument ```num_labels=1``` configures the final linear layer of the model to have a single output neuron which is appropriate for a regression task, as it leads to prediction of a single continuous value.\n\nThe first model was based on the [deberta-v3-small](https://huggingface.co/microsoft/deberta-v3-small) model ([link to article](https://openreview.net/forum?id=XPZIaotutsD)) with 44M parameters pre-trained on general text data. The \"sparse\" input used for the DeBERTa model was:\n```python\ntrdf['input'] = 'CELL TYPE: ' + trdf.cell_type + '; TREATMENT: ' + trdf.sm_name + '; GENE: ' + trdf.gene\n```\nFine-tuning the DeBERTa for 20-25 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 48-60 hours (see [Transformer DeBERTa (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo) for a training demo that runs on Kaggle in ~2 hours).\n\nThe second model was based on the [tiny-biobert](https://huggingface.co/nlpie/tiny-biobert) model ([link to article](https://doi.org/10.48550/arxiv.2209.03182)) with 15M parameters pre-trained on citations and abstracts of biomedical literature using the PubMed dataset. The \"verbose\" input used for the TinyBioBert model was:\n```python\ntrdf['input'] = \"Estimate the −log(p-value) confidence for change in gene expression of \" + trdf.gene + \" in \" + trdf.cell_type + \" when treated with \" + trdf.sm_name + \" compared to DMSO.\"\n```\nFine-tuning the TinyBioBert model for 20-30 epochs on an A10 GPU using an 80/20 train-valid-split and a batch size of 256 took about 24-30 hours (see [Transformer TinyBioBert (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo) for a training demo that runs on Kaggle in ~1:30 hours).\n\nInterestingly, the \"sparse\" input used for the DeBERTa model did not lead to very good performance in the TinyBioBert model, whereas the more \"verbose\" input used for the TinyBioBert model led to slower training and worse results in the DeBERTa model. This is reflected in the following submission scores (all trained for 5 epochs using the same random seed of 42):\n\n```python\n                        public LB     privateLB\nDeBERTa sparse:             0.630         0.807\nDeBERTa verbose:            0.637         0.810\nTinyBioBert sparse:         0.634         0.812\nTinyBioBert verbose:        0.624         0.807\n```\n\nThe transformer part of the final ensemble consisted of four TinyBioBert models. These were trained for 30 epochs using seed 42, 30 epochs using seed 546, 20 epochs using seed 419, and 5x5 epochs using different random seeds. Each model was assigned a weight of 0.15 in the Transformer sub-ensemble. The single DeBERTa model used was trained for a total of 22 epochs using different random seeds, and was assigned a weight of 0.4.\n## 4. Robustness\n\nThe robustness of each individual model type was evaluated using triplicate submissions with different random seeds by calculating the mean and standard deviation of the scores achieved on the public and private leaderboard (in the case of transformers, the models also varied slightly in the number of epochs trained). \n\nTabMod NN (600): 0.567 ± 0.003 public, 0.777 ± 0.007 private.\nTabMod NN (1000): 0.562 ± 0.002 public, 0.778 ± 0.003 private\n\nRandom Forest: 0.592 ± 0.004 public, 0.781 ± 0.002 private.\n\nXGBoosted Forest: 0.588 ± 0.002 public, 0.780 ± 0.001 private.\n\nTransformer (DeBERTa) 22-25 epochs: 0.606 ± 0.013 public, 0.784 ± 0.007 private.\nTransformer (TinyBioBert) 25-30 epochs: 0.607 ± 0.003 public, 0.777 ± 0.004 private.\n\nThese results show that the fine-tuned DeBERTa model is the least robust of the models. The fine-tuned TinyBioBert model consistently performed as well as the TabMod NNs on the private leaderboard, but consistently looked worse on the public leaderboard.\n## 5. Documentation & code style\n\nThe code is documented in the notebooks linked throughout the notebook. Additionally, a list of links to all notebooks and datasets can be found in Section 6.\n## 6. Reproducibility\n\nThe source code and further documentation for all models as well as additional datasets can be accessed through the following links:\n\nModel Notebooks:\n[TabMod NN Notebook](https://www.kaggle.com/code/frenio/30-op2scp-tabular-model-nn-with-pca10-denoising)\n[Random Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-random-forest-with-tabmod-embeds)\n[XGBoosted Forest Notebook](https://www.kaggle.com/code/frenio/30-op2scp-xgboosted-forest-with-tabmod-embeds)\n[Transformer DeBERTa (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-deberta-v3-small-demo)\n[Transformer TinyBioBert (Demo) Notebook](https://www.kaggle.com/code/frenio/30-op2scp-transformer-tinybiobert-demo)\n\nOther Notebooks:\n[Look at Embeddings Notebook](https://www.kaggle.com/code/frenio/30-op2scp-look-at-tabmod-nn-embeddings)\n[Prediction Correlation Notebook](https://www.kaggle.com/code/frenio/30-op2scp-correlation-of-predictions)\n\nDatasets:\n[OP2 Additional Features Dataset](https://www.kaggle.com/datasets/frenio/op2-scp-additional-cell-gene-and-mol-features)\n[OP2 TabMod NN Embeddings Dataset](https://www.kaggle.com/datasets/frenio/op2-single-cell-perturbations-tabmodnn-embeddings)",
    "2562757": "This comment used to contain Sections 3 to 6 of my solution write-up because a bug kept me from posting it in whole. Now that the bug is fixed, I replaced the original comment with this notice and added Sections 3 to 6 to the main topic.",
    "2569948": "Thanks for sharing, Frenio! Your score of 0.553 on the leaderboard really gave us a lot of nightmares. You led this challenge for a very long time!",
    "2562953": "Congratulations and thanks for sharing your precious insights (topic and comments above)",
    "2562826": "Thanks a lot for sharing ! \nAny ideas what can be interpretation of the two clusters for genes ? \nTo my mind there is no two natural clusters in human genes. There can be be different approaches giving different results , several are in HPA like: https://www.proteinatlas.org/humanproteome/tissue/expression+cluster\nPS\nWould you be so kind to make your key submit open , if possible ? \n",
    "2728295": "Hi, interesting and well described solution with melting the data and learning the embedding representations, thanks for sharing! In the Random Forest model, you used embeddings from a TabMod NN with a dimensionality of 1000 for genes and without using denoising on the training data, and as I understand it, it was better, while the TabMod NN itself performed better with embeddings after denoising and with a dimensionality of 600.\n\nMaybe it's just my poor understanding, but as I understand it, the better the model performs, the better the quality of the embeddings it learns during training. So how would you explain that lower quality embeddings performed better in RF?",
    "2579010": "Thanks again for the great write-up !\n\nI wonder about the \"Gene Information\" section - you mention that you convert genes information into features. \nWould you be so kind to clarify it a bit - because genes -are our targets  - so any info on genes -  do not depend on our features: cell type/drug, so it is a kind constant across sample-wise direction. There are enormous amounts   of information on human genes, but typically there is no way to use it in such kind of tasks, because it is \"constant\" across cells/cell types/drugs direction, so not a feature.\n",
    "2562754": ""
  }
}