{
  "id": 459573,
  "title": "4th Place Writeup",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/459573",
  "author_name": "",
  "post_date": "2023-12-05T20:32:01.617136500Z",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>Acknowledgements</strong><br>\nSecuring 4th place was an unexpected triumph, achieved with our final submission. We are excited about the success of our solution, and participating in this competition was both super fun and rewarding. We extend our gratitude to all competitors, especially those who generously shared their thoughts and work, as well as the hosts for organizing this engaging and challenging competition.</p>\n<p><strong>Context</strong><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Overview</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Data</a></p>\n<p><strong>Overview of the approach</strong><br>\nIn our approach, we conducted an extensive exploration of diverse model architectures and datasets, incorporating valuable chemical knowledge data. The final submission took the form of an ensemble, comprising a high-scoring Feedforward Neural Network ensemble consisting of 80 models (Public/Private LB of 0.571/0.765) and a complementary, lower-scoring, LGBM model (0.598/0.789). The Neural Networks demonstrated their prowess by capturing nuanced details from the training data, predicting all genes for each compound. In contrast, the LGBM model made decisions based on the differential expression (DE) of other cell types for each gene. For the Neural Network, we leveraged the mean predictions from 80 models to enhance the robustness of our ensemble. We also used one model Kishan our team lead shared publicly early on in the competition, based on NLP regression. The blending of these models was achieved through a weighted sum approach: 0.6 * W1(NN) + 0.2 * W2(LGBM) + 0.2*W3(NLP_Regression) and some post-processing, resulting in an ensemble score of (0.561/0.743).</p>\n<p>Notably, our success was further amplified by a strategic post-processing step applied in the final competition submission, adding a touch of magic to our winning solution and achieving an impressive score of (0.558/0.733).</p>\n<p>Our results can be reproduced pretty closely from notebooks we published and interconnected, for this writeup, see:</p>\n<p>We have 3 types of models:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net\" target=\"_blank\">Neural Network</a></li>\n<li><a href=\"https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation\" target=\"_blank\">LGBM</a></li>\n<li><a href=\"https://www.kaggle.com/code/kishanvavdara/nlp-regression\" target=\"_blank\">NLP</a></li>\n</ul>\n<p>That are ensembled in:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/raki21/4th-place-ensembling\" target=\"_blank\">Ensembling</a></li>\n</ul>\n<p>and finally passed to this:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\" target=\"_blank\">Postprocessing</a></li>\n</ul>\n<p><strong>Details of the submission</strong> </p>\n<p>Our winning submission had a touch of \"Last Day Magic\", which set it apart. On the final day, we discerned a pattern in low-variance compounds, realizing that for certain compound-gene combinations, the differential expression appeared more as random fluctuations than meaningful changes. In these instances, the model was fitting noise, capturing fluctuations rather than substantive shifts in gene expression. After confirming a strong correlation between predicted and true gene values, we swiftly implemented a last-minute post-processing adjustment. This tweak involved significantly shifting predictions for low-variance compounds towards zero, effectively filtering out the noise. Conceptually, it was akin to refining predictions for instances where the model was dominated by noise rather than meaningful signals.<br>\nFor instance:</p>\n<pre><code> compound_predictions  predictions:\n    If compound_predictions  small:\n         compound_predictions = \n</code></pre>\n<p>However, our approach was more nuanced, decreasing predictions significantly for low-variance compounds instead of outright setting them to zero. The degree of adjustment decreased for compounds with higher average predictions. After this you can increase magnitude of high variance compounds as previously prediction magnitude was too high for low variance compounds, it was too low for high variance compounds:</p>\n<pre><code>MULT = \n index, compound_gene_pred  pred.iterrows():\n    abs_compound_mean = (compound_gene_pred).mean()\n    compound_gene_pred *= (abs_compound_mean**, )\n    df.loc[index] = compound_gene_pred  \ndf[numeric_cols] *= MULT\n</code></pre>\n<p>The impact was remarkable, significantly boosting the scores of our two final submissions. One submission relied solely on our models, while the other maximally overfitted by incorporating a blend of blends tuned for each compound that circulated towards the end:</p>\n<p>Regular: 560/743 =&gt; 558/733<br>\nOverfit: 526/743 =&gt; 524/739</p>\n<p>These tweaks secured our 4th-place finish. You can explore a visualization and get the code to append it to your own solution in the <a href=\"https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\" target=\"_blank\">linked postprocessing notebook</a>, another user already <a href=\"https://www.kaggle.com/code/alexandervc/op2-explore-4th-place-magic?scriptVersionId=153724701&amp;cellId=17\" target=\"_blank\">verified that it works for many public submissions</a>! It's a reminder that even in the final hours of a competition, thoughtful analysis can make all the difference.</p>\n<p><strong>Feature selection and feature engineering</strong></p>\n<p>In refining feature selection for the Neural Network (NN), we experimented with diverse features, including SMILES fingerprint, SMILES as embedding, normalized cell counts, and cell_type correlation. Despite these efforts, none proved optimal. Our final successful feature set for NN comprised one-hot encoding of cell_type, sm_lincs_id, sm_name, and SMILES.</p>\n<p>For the Light Gradient Boosting Machine (LGBM), our feature selection encompassed compounds data from  <a href=\"https://lincsportal.ccs.miami.edu/dcic-portal/\" target=\"_blank\">LINCS</a> data portal and real cell count information sourced from <code>adata_obs_meta.csv</code> </p>\n<p><strong>Modelling</strong></p>\n<p>FeedForward Neural network </p>\n<p>Our FeedForward Neural Network (FNN) model is designed with a structured architecture that includes three one-hot encoded inputs and features eight hidden layers, resulting in an output dimension of 18211. The model underwent extensive training on the entire dataset, utilizing 80 different seeds and varying batch sizes between 32 and 64. To enhance model performance, the labels were normalized using a standard scaler. Through the amalgamation of mean predictions from various batches, our approach reached its pinnacle, resulting in a more resilient and adaptable model.</p>\n<p>LGBM</p>\n<p>The LGBM approach involves predicting not only segregated by compounds and cell type but also by gene. Initially, we focused on predicting differential expression (DE) for NK, T CD4+, and T regulatory cells, representing three inputs with one output. T CD8+ cells were excluded due to their low correlation with B/Myeloid cells, coupled with the absence of some compounds, which could have either reduced the train set size or required additional workarounds. Subsequently, we incorporated expression counts for each of the three donors for each cell type and gene, contributing to a total of +3*3 inputs. To further enrich our feature set, LINCS data was introduced as an additional input, bringing the total to +62 inputs. This comprehensive approach aimed to capture the nuances of gene expression across different cell types and compounds.<br>\nIn the post-processing phase, we aggregated gene information based on gene correlations specific to the cell type, utilizing a 18211x18211 correlation matrix. This matrix was multiplied with the outputs, considering gene relations, and divided by the magnitude of the correlation matrix row. While this process is challenging to visualize, you can find the implementation details in the public LGBM notebook.</p>\n<p>We experimented with various loss functions including Huber loss, Mean Squared Error (MSE), and a custom Mean Rowwise Root Mean Squared Error (MRRMSE), but none outperformed Mean Absolute Error (MAE). Consequently, we adhered to using MAE as our preferred loss function for its superior performance in our NN model.</p>\n<p>In the case of LGBM, RMSE and MAE exhibited comparable performance, with a slight edge for RMSE. </p>\n<p><strong>Validation strategy</strong></p>\n<p>We employed two validation strategies with different seeds for our models. Specifically, for the Neural Network, we utilized k-fold validation with MRRMSE. In the case of LGBM, our strategy involved validation on a per-compound basis, utilizing group k-fold on the 15 non-control train samples. To gain additional insights into our model's performance, we visualized the predictions. </p>\n<p>We placed significant emphasis on public Leaderboard scores due to the larger sample size for B and Myeloid compared to the training set.  As we trained our initial robust models, visualizations of prediction magnitudes on train cross-validation, public, and private test datasets revealed that the public test exhibited more similarities to the private test than the training set. This observation guided our strategy and reinforced the importance of leveraging public Leaderboard scores for better generalization.</p>\n<p><strong>Preventing Overfitting</strong></p>\n<p>Addressing the constraints posed by the limited dataset size, a key emphasis was placed on implementing effective regularization to prevent overfitting. To achieve this, a critical measure involved introducing a dropout layer with a rate of 0.5 before the output layer in the Neural Network (NN). Additionally, overfitting was mitigated by adopting an early stopping strategy. Specifically, model training was halted once the (MRRMSE) dropped below 0.5999 on the full dataset during training, providing a proactive approach to prevent overfitting.</p>\n<p>Overfitting in the LGBM model was primarily mitigated through early stopping, while the inclusion of L1 and L2 loss had minimal impact on addressing overfitting concerns.</p>\n<p><strong>Sources</strong><br>\n<a href=\"https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net\" target=\"_blank\">https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net</a><br>\n<a href=\"https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation\" target=\"_blank\">https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation</a><br>\n<a href=\"https://www.kaggle.com/code/kishanvavdara/nlp-regression\" target=\"_blank\">https://www.kaggle.com/code/kishanvavdara/nlp-regression</a><br>\n<a href=\"https://www.kaggle.com/raki21/4th-place-ensembling\" target=\"_blank\">https://www.kaggle.com/raki21/4th-place-ensembling</a><br>\n<a href=\"https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\" target=\"_blank\">https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing</a></p>\n<p><a href=\"https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda\" target=\"_blank\">https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda</a><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s</a></p>",
  "messages": [
    {
      "id": "2550148",
      "postDate": "12/05/2023 20:32:01",
      "content": "<p><strong>Acknowledgements</strong><br>\nSecuring 4th place was an unexpected triumph, achieved with our final submission. We are excited about the success of our solution, and participating in this competition was both super fun and rewarding. We extend our gratitude to all competitors, especially those who generously shared their thoughts and work, as well as the hosts for organizing this engaging and challenging competition.</p>\n<p><strong>Context</strong><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Overview</a><br>\n<a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Data</a></p>\n<p><strong>Overview of the approach</strong><br>\nIn our approach, we conducted an extensive exploration of diverse model architectures and datasets, incorporating valuable chemical knowledge data. The final submission took the form of an ensemble, comprising a high-scoring Feedforward Neural Network ensemble consisting of 80 models (Public/Private LB of 0.571/0.765) and a complementary, lower-scoring, LGBM model (0.598/0.789). The Neural Networks demonstrated their prowess by capturing nuanced details from the training data, predicting all genes for each compound. In contrast, the LGBM model made decisions based on the differential expression (DE) of other cell types for each gene. For the Neural Network, we leveraged the mean predictions from 80 models to enhance the robustness of our ensemble. We also used one model Kishan our team lead shared publicly early on in the competition, based on NLP regression. The blending of these models was achieved through a weighted sum approach: 0.6 * W1(NN) + 0.2 * W2(LGBM) + 0.2*W3(NLP_Regression) and some post-processing, resulting in an ensemble score of (0.561/0.743).</p>\n<p>Notably, our success was further amplified by a strategic post-processing step applied in the final competition submission, adding a touch of magic to our winning solution and achieving an impressive score of (0.558/0.733).</p>\n<p>Our results can be reproduced pretty closely from notebooks we published and interconnected, for this writeup, see:</p>\n<p>We have 3 types of models:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net\" target=\"_blank\">Neural Network</a></li>\n<li><a href=\"https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation\" target=\"_blank\">LGBM</a></li>\n<li><a href=\"https://www.kaggle.com/code/kishanvavdara/nlp-regression\" target=\"_blank\">NLP</a></li>\n</ul>\n<p>That are ensembled in:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/raki21/4th-place-ensembling\" target=\"_blank\">Ensembling</a></li>\n</ul>\n<p>and finally passed to this:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\" target=\"_blank\">Postprocessing</a></li>\n</ul>\n<p><strong>Details of the submission</strong> </p>\n<p>Our winning submission had a touch of \"Last Day Magic\", which set it apart. On the final day, we discerned a pattern in low-variance compounds, realizing that for certain compound-gene combinations, the differential expression appeared more as random fluctuations than meaningful changes. In these instances, the model was fitting noise, capturing fluctuations rather than substantive shifts in gene expression. After confirming a strong correlation between predicted and true gene values, we swiftly implemented a last-minute post-processing adjustment. This tweak involved significantly shifting predictions for low-variance compounds towards zero, effectively filtering out the noise. Conceptually, it was akin to refining predictions for instances where the model was dominated by noise rather than meaningful signals.<br>\nFor instance:</p>\n<pre><code> compound_predictions  predictions:\n    If compound_predictions  small:\n         compound_predictions = \n</code></pre>\n<p>However, our approach was more nuanced, decreasing predictions significantly for low-variance compounds instead of outright setting them to zero. The degree of adjustment decreased for compounds with higher average predictions. After this you can increase magnitude of high variance compounds as previously prediction magnitude was too high for low variance compounds, it was too low for high variance compounds:</p>\n<pre><code>MULT = \n index, compound_gene_pred  pred.iterrows():\n    abs_compound_mean = (compound_gene_pred).mean()\n    compound_gene_pred *= (abs_compound_mean**, )\n    df.loc[index] = compound_gene_pred  \ndf[numeric_cols] *= MULT\n</code></pre>\n<p>The impact was remarkable, significantly boosting the scores of our two final submissions. One submission relied solely on our models, while the other maximally overfitted by incorporating a blend of blends tuned for each compound that circulated towards the end:</p>\n<p>Regular: 560/743 =&gt; 558/733<br>\nOverfit: 526/743 =&gt; 524/739</p>\n<p>These tweaks secured our 4th-place finish. You can explore a visualization and get the code to append it to your own solution in the <a href=\"https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\" target=\"_blank\">linked postprocessing notebook</a>, another user already <a href=\"https://www.kaggle.com/code/alexandervc/op2-explore-4th-place-magic?scriptVersionId=153724701&amp;cellId=17\" target=\"_blank\">verified that it works for many public submissions</a>! It's a reminder that even in the final hours of a competition, thoughtful analysis can make all the difference.</p>\n<p><strong>Feature selection and feature engineering</strong></p>\n<p>In refining feature selection for the Neural Network (NN), we experimented with diverse features, including SMILES fingerprint, SMILES as embedding, normalized cell counts, and cell_type correlation. Despite these efforts, none proved optimal. Our final successful feature set for NN comprised one-hot encoding of cell_type, sm_lincs_id, sm_name, and SMILES.</p>\n<p>For the Light Gradient Boosting Machine (LGBM), our feature selection encompassed compounds data from  <a href=\"https://lincsportal.ccs.miami.edu/dcic-portal/\" target=\"_blank\">LINCS</a> data portal and real cell count information sourced from <code>adata_obs_meta.csv</code> </p>\n<p><strong>Modelling</strong></p>\n<p>FeedForward Neural network </p>\n<p>Our FeedForward Neural Network (FNN) model is designed with a structured architecture that includes three one-hot encoded inputs and features eight hidden layers, resulting in an output dimension of 18211. The model underwent extensive training on the entire dataset, utilizing 80 different seeds and varying batch sizes between 32 and 64. To enhance model performance, the labels were normalized using a standard scaler. Through the amalgamation of mean predictions from various batches, our approach reached its pinnacle, resulting in a more resilient and adaptable model.</p>\n<p>LGBM</p>\n<p>The LGBM approach involves predicting not only segregated by compounds and cell type but also by gene. Initially, we focused on predicting differential expression (DE) for NK, T CD4+, and T regulatory cells, representing three inputs with one output. T CD8+ cells were excluded due to their low correlation with B/Myeloid cells, coupled with the absence of some compounds, which could have either reduced the train set size or required additional workarounds. Subsequently, we incorporated expression counts for each of the three donors for each cell type and gene, contributing to a total of +3*3 inputs. To further enrich our feature set, LINCS data was introduced as an additional input, bringing the total to +62 inputs. This comprehensive approach aimed to capture the nuances of gene expression across different cell types and compounds.<br>\nIn the post-processing phase, we aggregated gene information based on gene correlations specific to the cell type, utilizing a 18211x18211 correlation matrix. This matrix was multiplied with the outputs, considering gene relations, and divided by the magnitude of the correlation matrix row. While this process is challenging to visualize, you can find the implementation details in the public LGBM notebook.</p>\n<p>We experimented with various loss functions including Huber loss, Mean Squared Error (MSE), and a custom Mean Rowwise Root Mean Squared Error (MRRMSE), but none outperformed Mean Absolute Error (MAE). Consequently, we adhered to using MAE as our preferred loss function for its superior performance in our NN model.</p>\n<p>In the case of LGBM, RMSE and MAE exhibited comparable performance, with a slight edge for RMSE. </p>\n<p><strong>Validation strategy</strong></p>\n<p>We employed two validation strategies with different seeds for our models. Specifically, for the Neural Network, we utilized k-fold validation with MRRMSE. In the case of LGBM, our strategy involved validation on a per-compound basis, utilizing group k-fold on the 15 non-control train samples. To gain additional insights into our model's performance, we visualized the predictions. </p>\n<p>We placed significant emphasis on public Leaderboard scores due to the larger sample size for B and Myeloid compared to the training set.  As we trained our initial robust models, visualizations of prediction magnitudes on train cross-validation, public, and private test datasets revealed that the public test exhibited more similarities to the private test than the training set. This observation guided our strategy and reinforced the importance of leveraging public Leaderboard scores for better generalization.</p>\n<p><strong>Preventing Overfitting</strong></p>\n<p>Addressing the constraints posed by the limited dataset size, a key emphasis was placed on implementing effective regularization to prevent overfitting. To achieve this, a critical measure involved introducing a dropout layer with a rate of 0.5 before the output layer in the Neural Network (NN). Additionally, overfitting was mitigated by adopting an early stopping strategy. Specifically, model training was halted once the (MRRMSE) dropped below 0.5999 on the full dataset during training, providing a proactive approach to prevent overfitting.</p>\n<p>Overfitting in the LGBM model was primarily mitigated through early stopping, while the inclusion of L1 and L2 loss had minimal impact on addressing overfitting concerns.</p>\n<p><strong>Sources</strong><br>\n<a href=\"https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net\" target=\"_blank\">https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net</a><br>\n<a href=\"https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation\" target=\"_blank\">https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation</a><br>\n<a href=\"https://www.kaggle.com/code/kishanvavdara/nlp-regression\" target=\"_blank\">https://www.kaggle.com/code/kishanvavdara/nlp-regression</a><br>\n<a href=\"https://www.kaggle.com/raki21/4th-place-ensembling\" target=\"_blank\">https://www.kaggle.com/raki21/4th-place-ensembling</a><br>\n<a href=\"https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\" target=\"_blank\">https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing</a></p>\n<p><a href=\"https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda\" target=\"_blank\">https://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda</a><br>\n<a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s</a></p>",
      "rawMarkdown": "**Acknowledgements**\nSecuring 4th place was an unexpected triumph, achieved with our final submission. We are excited about the success of our solution, and participating in this competition was both super fun and rewarding. We extend our gratitude to all competitors, especially those who generously shared their thoughts and work, as well as the hosts for organizing this engaging and challenging competition.\n\n**Context**\n[Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n[Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n**Overview of the approach**\nIn our approach, we conducted an extensive exploration of diverse model architectures and datasets, incorporating valuable chemical knowledge data. The final submission took the form of an ensemble, comprising a high-scoring Feedforward Neural Network ensemble consisting of 80 models (Public/Private LB of 0.571/0.765) and a complementary, lower-scoring, LGBM model (0.598/0.789). The Neural Networks demonstrated their prowess by capturing nuanced details from the training data, predicting all genes for each compound. In contrast, the LGBM model made decisions based on the differential expression (DE) of other cell types for each gene. For the Neural Network, we leveraged the mean predictions from 80 models to enhance the robustness of our ensemble. We also used one model Kishan our team lead shared publicly early on in the competition, based on NLP regression. The blending of these models was achieved through a weighted sum approach: 0.6 * W1(NN) + 0.2 * W2(LGBM) + 0.2*W3(NLP_Regression) and some post-processing, resulting in an ensemble score of (0.561/0.743).\n\nNotably, our success was further amplified by a strategic post-processing step applied in the final competition submission, adding a touch of magic to our winning solution and achieving an impressive score of (0.558/0.733).\n\nOur results can be reproduced pretty closely from notebooks we published and interconnected, for this writeup, see:\n\nWe have 3 types of models:\n* [Neural Network](https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net)\n* [LGBM](https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation)\n* [NLP](https://www.kaggle.com/code/kishanvavdara/nlp-regression)\n\nThat are ensembled in:\n* [Ensembling](https://www.kaggle.com/raki21/4th-place-ensembling)\n\nand finally passed to this:\n* [Postprocessing](https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing)\n\n**Details of the submission** \n\nOur winning submission had a touch of \"Last Day Magic\", which set it apart. On the final day, we discerned a pattern in low-variance compounds, realizing that for certain compound-gene combinations, the differential expression appeared more as random fluctuations than meaningful changes. In these instances, the model was fitting noise, capturing fluctuations rather than substantive shifts in gene expression. After confirming a strong correlation between predicted and true gene values, we swiftly implemented a last-minute post-processing adjustment. This tweak involved significantly shifting predictions for low-variance compounds towards zero, effectively filtering out the noise. Conceptually, it was akin to refining predictions for instances where the model was dominated by noise rather than meaningful signals.\nFor instance:\n\n```python\nfor compound_predictions in predictions:\n\tIf compound_predictions is small:\n\t\tSet compound_predictions = 0\n```\n\nHowever, our approach was more nuanced, decreasing predictions significantly for low-variance compounds instead of outright setting them to zero. The degree of adjustment decreased for compounds with higher average predictions. After this you can increase magnitude of high variance compounds as previously prediction magnitude was too high for low variance compounds, it was too low for high variance compounds:\n\n```python\nMULT = 1.3\nfor index, compound_gene_pred in pred.iterrows():\n    abs_compound_mean = abs(compound_gene_pred).mean()\n    compound_gene_pred *= min(abs_compound_mean**0.6, 1)\n    df.loc[index] = compound_gene_pred  \ndf[numeric_cols] *= MULT\n```\n\nThe impact was remarkable, significantly boosting the scores of our two final submissions. One submission relied solely on our models, while the other maximally overfitted by incorporating a blend of blends tuned for each compound that circulated towards the end:\n\nRegular: 560/743 => 558/733\nOverfit: 526/743 => 524/739\n\nThese tweaks secured our 4th-place finish. You can explore a visualization and get the code to append it to your own solution in the [linked postprocessing notebook](https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing), another user already [verified that it works for many public submissions](https://www.kaggle.com/code/alexandervc/op2-explore-4th-place-magic?scriptVersionId=153724701&cellId=17)! It's a reminder that even in the final hours of a competition, thoughtful analysis can make all the difference.\n\n**Feature selection and feature engineering**\n\nIn refining feature selection for the Neural Network (NN), we experimented with diverse features, including SMILES fingerprint, SMILES as embedding, normalized cell counts, and cell_type correlation. Despite these efforts, none proved optimal. Our final successful feature set for NN comprised one-hot encoding of cell_type, sm_lincs_id, sm_name, and SMILES.\n\nFor the Light Gradient Boosting Machine (LGBM), our feature selection encompassed compounds data from  [LINCS](https://lincsportal.ccs.miami.edu/dcic-portal/) data portal and real cell count information sourced from `adata_obs_meta.csv` \n\n\n**Modelling**\n \nFeedForward Neural network \n\nOur FeedForward Neural Network (FNN) model is designed with a structured architecture that includes three one-hot encoded inputs and features eight hidden layers, resulting in an output dimension of 18211. The model underwent extensive training on the entire dataset, utilizing 80 different seeds and varying batch sizes between 32 and 64. To enhance model performance, the labels were normalized using a standard scaler. Through the amalgamation of mean predictions from various batches, our approach reached its pinnacle, resulting in a more resilient and adaptable model.\n\n  \n\t\nLGBM\n\t\nThe LGBM approach involves predicting not only segregated by compounds and cell type but also by gene. Initially, we focused on predicting differential expression (DE) for NK, T CD4+, and T regulatory cells, representing three inputs with one output. T CD8+ cells were excluded due to their low correlation with B/Myeloid cells, coupled with the absence of some compounds, which could have either reduced the train set size or required additional workarounds. Subsequently, we incorporated expression counts for each of the three donors for each cell type and gene, contributing to a total of +3*3 inputs. To further enrich our feature set, LINCS data was introduced as an additional input, bringing the total to +62 inputs. This comprehensive approach aimed to capture the nuances of gene expression across different cell types and compounds.\nIn the post-processing phase, we aggregated gene information based on gene correlations specific to the cell type, utilizing a 18211x18211 correlation matrix. This matrix was multiplied with the outputs, considering gene relations, and divided by the magnitude of the correlation matrix row. While this process is challenging to visualize, you can find the implementation details in the public LGBM notebook.\n\nWe experimented with various loss functions including Huber loss, Mean Squared Error (MSE), and a custom Mean Rowwise Root Mean Squared Error (MRRMSE), but none outperformed Mean Absolute Error (MAE). Consequently, we adhered to using MAE as our preferred loss function for its superior performance in our NN model.\n\nIn the case of LGBM, RMSE and MAE exhibited comparable performance, with a slight edge for RMSE. \n\n\t\n**Validation strategy**\n\nWe employed two validation strategies with different seeds for our models. Specifically, for the Neural Network, we utilized k-fold validation with MRRMSE. In the case of LGBM, our strategy involved validation on a per-compound basis, utilizing group k-fold on the 15 non-control train samples. To gain additional insights into our model's performance, we visualized the predictions. \n\nWe placed significant emphasis on public Leaderboard scores due to the larger sample size for B and Myeloid compared to the training set.  As we trained our initial robust models, visualizations of prediction magnitudes on train cross-validation, public, and private test datasets revealed that the public test exhibited more similarities to the private test than the training set. This observation guided our strategy and reinforced the importance of leveraging public Leaderboard scores for better generalization.\n\n**Preventing Overfitting**\n\nAddressing the constraints posed by the limited dataset size, a key emphasis was placed on implementing effective regularization to prevent overfitting. To achieve this, a critical measure involved introducing a dropout layer with a rate of 0.5 before the output layer in the Neural Network (NN). Additionally, overfitting was mitigated by adopting an early stopping strategy. Specifically, model training was halted once the (MRRMSE) dropped below 0.5999 on the full dataset during training, providing a proactive approach to prevent overfitting.\n\nOverfitting in the LGBM model was primarily mitigated through early stopping, while the inclusion of L1 and L2 loss had minimal impact on addressing overfitting concerns.\n\n\n**Sources**\nhttps://www.kaggle.com/code/kishanvavdara/4th-place-neural-net\nhttps://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation\nhttps://www.kaggle.com/code/kishanvavdara/nlp-regression\nhttps://www.kaggle.com/raki21/4th-place-ensembling\nhttps://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\n\nhttps://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s",
      "votes": null
    },
    {
      "id": "2550157",
      "postDate": "12/05/2023 20:37:51",
      "content": "<p>We saw that we were disqualified an hour ago even though we are not aware of any rule violations, so we hope this clarifies that we worked honestly on this challenge and can account for how our results came about. We used all channels we know of to provide Kaggle with information and hope we can resolve this as soon as possible. We can currently not set this as a linked writeup though, but would do this as soon as we are back on the leaderboard.</p>",
      "rawMarkdown": "We saw that we were disqualified an hour ago even though we are not aware of any rule violations, so we hope this clarifies that we worked honestly on this challenge and can account for how our results came about. We used all channels we know of to provide Kaggle with information and hope we can resolve this as soon as possible. We can currently not set this as a linked writeup though, but would do this as soon as we are back on the leaderboard.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2550157,
      "author_name": "raki21",
      "author_url": "",
      "post_date": "12/05/2023 20:37:51",
      "content": "<p>We saw that we were disqualified an hour ago even though we are not aware of any rule violations, so we hope this clarifies that we worked honestly on this challenge and can account for how our results came about. We used all channels we know of to provide Kaggle with information and hope we can resolve this as soon as possible. We can currently not set this as a linked writeup though, but would do this as soon as we are back on the leaderboard.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2550148": "**Acknowledgements**\nSecuring 4th place was an unexpected triumph, achieved with our final submission. We are excited about the success of our solution, and participating in this competition was both super fun and rewarding. We extend our gratitude to all competitors, especially those who generously shared their thoughts and work, as well as the hosts for organizing this engaging and challenging competition.\n\n**Context**\n[Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n[Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n**Overview of the approach**\nIn our approach, we conducted an extensive exploration of diverse model architectures and datasets, incorporating valuable chemical knowledge data. The final submission took the form of an ensemble, comprising a high-scoring Feedforward Neural Network ensemble consisting of 80 models (Public/Private LB of 0.571/0.765) and a complementary, lower-scoring, LGBM model (0.598/0.789). The Neural Networks demonstrated their prowess by capturing nuanced details from the training data, predicting all genes for each compound. In contrast, the LGBM model made decisions based on the differential expression (DE) of other cell types for each gene. For the Neural Network, we leveraged the mean predictions from 80 models to enhance the robustness of our ensemble. We also used one model Kishan our team lead shared publicly early on in the competition, based on NLP regression. The blending of these models was achieved through a weighted sum approach: 0.6 * W1(NN) + 0.2 * W2(LGBM) + 0.2*W3(NLP_Regression) and some post-processing, resulting in an ensemble score of (0.561/0.743).\n\nNotably, our success was further amplified by a strategic post-processing step applied in the final competition submission, adding a touch of magic to our winning solution and achieving an impressive score of (0.558/0.733).\n\nOur results can be reproduced pretty closely from notebooks we published and interconnected, for this writeup, see:\n\nWe have 3 types of models:\n* [Neural Network](https://www.kaggle.com/code/kishanvavdara/4th-place-neural-net)\n* [LGBM](https://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation)\n* [NLP](https://www.kaggle.com/code/kishanvavdara/nlp-regression)\n\nThat are ensembled in:\n* [Ensembling](https://www.kaggle.com/raki21/4th-place-ensembling)\n\nand finally passed to this:\n* [Postprocessing](https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing)\n\n**Details of the submission** \n\nOur winning submission had a touch of \"Last Day Magic\", which set it apart. On the final day, we discerned a pattern in low-variance compounds, realizing that for certain compound-gene combinations, the differential expression appeared more as random fluctuations than meaningful changes. In these instances, the model was fitting noise, capturing fluctuations rather than substantive shifts in gene expression. After confirming a strong correlation between predicted and true gene values, we swiftly implemented a last-minute post-processing adjustment. This tweak involved significantly shifting predictions for low-variance compounds towards zero, effectively filtering out the noise. Conceptually, it was akin to refining predictions for instances where the model was dominated by noise rather than meaningful signals.\nFor instance:\n\n```python\nfor compound_predictions in predictions:\n\tIf compound_predictions is small:\n\t\tSet compound_predictions = 0\n```\n\nHowever, our approach was more nuanced, decreasing predictions significantly for low-variance compounds instead of outright setting them to zero. The degree of adjustment decreased for compounds with higher average predictions. After this you can increase magnitude of high variance compounds as previously prediction magnitude was too high for low variance compounds, it was too low for high variance compounds:\n\n```python\nMULT = 1.3\nfor index, compound_gene_pred in pred.iterrows():\n    abs_compound_mean = abs(compound_gene_pred).mean()\n    compound_gene_pred *= min(abs_compound_mean**0.6, 1)\n    df.loc[index] = compound_gene_pred  \ndf[numeric_cols] *= MULT\n```\n\nThe impact was remarkable, significantly boosting the scores of our two final submissions. One submission relied solely on our models, while the other maximally overfitted by incorporating a blend of blends tuned for each compound that circulated towards the end:\n\nRegular: 560/743 => 558/733\nOverfit: 526/743 => 524/739\n\nThese tweaks secured our 4th-place finish. You can explore a visualization and get the code to append it to your own solution in the [linked postprocessing notebook](https://www.kaggle.com/code/raki21/4th-place-magic-postprocessing), another user already [verified that it works for many public submissions](https://www.kaggle.com/code/alexandervc/op2-explore-4th-place-magic?scriptVersionId=153724701&cellId=17)! It's a reminder that even in the final hours of a competition, thoughtful analysis can make all the difference.\n\n**Feature selection and feature engineering**\n\nIn refining feature selection for the Neural Network (NN), we experimented with diverse features, including SMILES fingerprint, SMILES as embedding, normalized cell counts, and cell_type correlation. Despite these efforts, none proved optimal. Our final successful feature set for NN comprised one-hot encoding of cell_type, sm_lincs_id, sm_name, and SMILES.\n\nFor the Light Gradient Boosting Machine (LGBM), our feature selection encompassed compounds data from  [LINCS](https://lincsportal.ccs.miami.edu/dcic-portal/) data portal and real cell count information sourced from `adata_obs_meta.csv` \n\n\n**Modelling**\n \nFeedForward Neural network \n\nOur FeedForward Neural Network (FNN) model is designed with a structured architecture that includes three one-hot encoded inputs and features eight hidden layers, resulting in an output dimension of 18211. The model underwent extensive training on the entire dataset, utilizing 80 different seeds and varying batch sizes between 32 and 64. To enhance model performance, the labels were normalized using a standard scaler. Through the amalgamation of mean predictions from various batches, our approach reached its pinnacle, resulting in a more resilient and adaptable model.\n\n  \n\t\nLGBM\n\t\nThe LGBM approach involves predicting not only segregated by compounds and cell type but also by gene. Initially, we focused on predicting differential expression (DE) for NK, T CD4+, and T regulatory cells, representing three inputs with one output. T CD8+ cells were excluded due to their low correlation with B/Myeloid cells, coupled with the absence of some compounds, which could have either reduced the train set size or required additional workarounds. Subsequently, we incorporated expression counts for each of the three donors for each cell type and gene, contributing to a total of +3*3 inputs. To further enrich our feature set, LINCS data was introduced as an additional input, bringing the total to +62 inputs. This comprehensive approach aimed to capture the nuances of gene expression across different cell types and compounds.\nIn the post-processing phase, we aggregated gene information based on gene correlations specific to the cell type, utilizing a 18211x18211 correlation matrix. This matrix was multiplied with the outputs, considering gene relations, and divided by the magnitude of the correlation matrix row. While this process is challenging to visualize, you can find the implementation details in the public LGBM notebook.\n\nWe experimented with various loss functions including Huber loss, Mean Squared Error (MSE), and a custom Mean Rowwise Root Mean Squared Error (MRRMSE), but none outperformed Mean Absolute Error (MAE). Consequently, we adhered to using MAE as our preferred loss function for its superior performance in our NN model.\n\nIn the case of LGBM, RMSE and MAE exhibited comparable performance, with a slight edge for RMSE. \n\n\t\n**Validation strategy**\n\nWe employed two validation strategies with different seeds for our models. Specifically, for the Neural Network, we utilized k-fold validation with MRRMSE. In the case of LGBM, our strategy involved validation on a per-compound basis, utilizing group k-fold on the 15 non-control train samples. To gain additional insights into our model's performance, we visualized the predictions. \n\nWe placed significant emphasis on public Leaderboard scores due to the larger sample size for B and Myeloid compared to the training set.  As we trained our initial robust models, visualizations of prediction magnitudes on train cross-validation, public, and private test datasets revealed that the public test exhibited more similarities to the private test than the training set. This observation guided our strategy and reinforced the importance of leveraging public Leaderboard scores for better generalization.\n\n**Preventing Overfitting**\n\nAddressing the constraints posed by the limited dataset size, a key emphasis was placed on implementing effective regularization to prevent overfitting. To achieve this, a critical measure involved introducing a dropout layer with a rate of 0.5 before the output layer in the Neural Network (NN). Additionally, overfitting was mitigated by adopting an early stopping strategy. Specifically, model training was halted once the (MRRMSE) dropped below 0.5999 on the full dataset during training, providing a proactive approach to prevent overfitting.\n\nOverfitting in the LGBM model was primarily mitigated through early stopping, while the inclusion of L1 and L2 loss had minimal impact on addressing overfitting concerns.\n\n\n**Sources**\nhttps://www.kaggle.com/code/kishanvavdara/4th-place-neural-net\nhttps://www.kaggle.com/raki21/4th-place-lgbm-with-gene-aggregation\nhttps://www.kaggle.com/code/kishanvavdara/nlp-regression\nhttps://www.kaggle.com/raki21/4th-place-ensembling\nhttps://www.kaggle.com/code/raki21/4th-place-magic-postprocessing\n\nhttps://www.kaggle.com/code/awater1223/op2-00-basic-metadata-eda\nhttps://www.kaggle.com/code/alexandervc/op2-eda-baseline-s",
    "2550157": "We saw that we were disqualified an hour ago even though we are not aware of any rule violations, so we hope this clarifies that we worked honestly on this challenge and can account for how our results came about. We used all channels we know of to provide Kaggle with information and hope we can resolve this as soon as possible. We can currently not set this as a linked writeup though, but would do this as soon as we are back on the leaderboard."
  },
  "source": "meta"
}