{
  "id": 458738,
  "title": "2nd Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/458738",
  "author_name": "EliKal",
  "post_date": "2023-12-01T09:34:36.335000",
  "votes": 49,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I am thrilled to finally unveil my solution to this competition!</p>\n<h2>Context</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>My optimal model emerged as a composite of four models, where three were employed as a normalized sum, and the fourth acted as an amplifier, showcasing discernible impact, potentially on labels with alternating signs:</p>\n<ul>\n<li><p>weight_df1: 0.5 (utilizing std, mean, and clustering sampling, yielding 0.551)</p></li>\n<li><p>weight_df2: 0.25 (excluding uncommon elements, resulting in 0.559)</p></li>\n<li><p>weight_df3: 0.25 (leveraging clustering sampling, achieving 0.575)</p></li>\n<li><p>weight_df4: 0.3 (incorporating mean, random sampling, and excluding std, attaining 0.554)<br>\nThe resultant model is expressed as the weighted sum of these components: resulted_model = weight_df1 * df1 + weight_df2 * df2 + weight_df3 * df3 + weight_df4 * df4. All models were trained using the same transformer architecture, albeit with varying dimensional models (with the optimal dimension being 128).</p></li>\n</ul>\n<p><strong>Data preprocessing and feature selection</strong><br>\nTo train deep learning models on categorical labels, I employed one-hot encoding transformation. To address high and low bias labels, I utilized target encoding by calculating the mean and standard deviation for each cell type and SM name. During experimentation, I compared models with and without uncommon columns and discovered that including uncommon columns significantly improved the training of encoding layers for mean and standard deviation feature vectors, resulting in noticeable performance enhancement.<br>\n<strong>EDA</strong></p>\n<p>In my feature exploration, I performed truncated Singular Value Decomposition (SVD) specifically on the target variables. Despite experimenting with different SVD sizes, the models generated from this approach consistently fell short of replicating the performance observed in the full targets regression. Notably, as part of this analysis, I identified certain targets with significant standard deviation values.</p>\n<p>To leverage this insight, I strategically enriched the feature set by introducing the standard deviation as a dedicated feature for each cell type and SM name. This involved concatenating these standard deviation values into a unified vector format (std_cell_type, std_sm_name).</p>\n<p>Furthermore, as I delved into the dataset, I conducted a comprehensive examination, revealing a remarkable cleanliness characterized by the absence of NaN values and duplicates. This meticulous exploration not only extended to model training but also encompassed a nuanced analysis of the target variables, contributing to the informed enhancement of the feature set for improved model performance.</p>\n<p><strong>Sampling strategy</strong><br>\nIn devising an effective sampling strategy for partitioning the data into training and validation sets, I employed a sophisticated approach rooted in the clusters derived from a K-Means clustering analysis on the target values. The rationale behind this strategy lies in the premise that data points within similar clusters share inherent patterns or characteristics, thus ensuring a more nuanced and representative split.</p>\n<p>For each distinct cluster identified through K-Means, I executed the data partitioning using the train_test_split function from the scikit-learn library. This method facilitated a controlled allocation of data points to both the training and validation sets, ensuring that the models were exposed to a diverse yet stratified representation of the data.</p>\n<p>The validation percentage, a crucial parameter in this process, was meticulously chosen to strike a balance between model robustness and computational efficiency. By experimenting with validation percentages ranging from 0.1 to 0.2, I systematically assessed the impact on model performance. Ultimately, for the optimal model configuration, a validation percentage of 0.1 was deemed most effective. This decision was guided by a careful consideration of the trade-off between the need for a sufficiently large training set and the importance of a robust validation set to gauge model generalizability.</p>\n<p>In summary, this sampling strategy, grounded in cluster-based data splitting, not only acknowledges the underlying structure within the target values but also optimizes the partitioning process to enhance the model's ability to capture diverse patterns and generalize effectively to unseen data.</p>\n<p><strong>Modeling</strong></p>\n<p>This is my best architecture</p>\n<pre><code>  (nn.Module):  \n     ():\n        (CustomTransformer_v3, self).__init__()\n        self.num_target_encodings =  * \n        self.num_sparse_features = num_features - self.num_target_encodings\n\n        self.sparse_feature_embedding = nn.Linear(self.num_sparse_features, d_model)\n        self.target_encoding_embedding = nn.Linear(self.num_target_encodings, d_model)\n        self.norm = nn.LayerNorm(d_model)\n\n        self.concatenation_layer = nn.Linear( * d_model, d_model)\n        self.transformer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=num_heads, dropout=dropout, activation=nn.GELU(),\n                                       batch_first=),\n            num_layers=num_layers\n        )\n        self.fc = nn.Linear(d_model, num_labels)\n\n     ():\n        sparse_features = x[:, :self.num_sparse_features]\n        target_encodings = x[:, self.num_sparse_features:]\n\n        sparse_features = self.sparse_feature_embedding(sparse_features)\n        target_encodings = self.target_encoding_embedding(target_encodings)\n\n        combined_features = torch.cat((sparse_features, target_encodings), dim=)\n        combined_features = self.concatenation_layer(combined_features)\n        combined_features = self.norm(combined_features)\n\n        x = self.transformer(combined_features)\n        x = self.norm(x)\n\n        x = self.fc(x)\n         x\n</code></pre>\n<p>I modularized my model into two distinct segments to efficiently handle the complexity of the input data:</p>\n<p>Sparse Feature Encoding:</p>\n<p>For encoding sparse features, I utilized an embedding layer tailored to handle the sparsity inherent in certain feature types.<br>\nConsidering the nature of sparse features, I initially explored utilizing the nn.Embedding layer. However, due to computational constraints on my laptop GPU, I opted for an alternative approach using a linear layer (nn.Linear) to convert sparse features into a dense representation.<br>\nTarget Encoding Feature Encoding (Dense Features):</p>\n<p>Concurrently, I addressed the encoding of target encodings, which are inherently dense features.<br>\nTo accomplish this, I employed a separate linear layer (nn.Linear) designed specifically for target encodings, ensuring an effective transformation into a meaningful latent space.<br>\nFollowing these individual encodings, I concatenated the resulting embeddings into a unified latent vector, fostering a comprehensive representation of both sparse and dense feature information. To further enhance the model's capacity to capture nuanced patterns, I employed normalization (nn.LayerNorm) on the concatenated features.</p>\n<p>Considering the computational challenges posed by sparse feature embedding using nn.Embedding, I strategically utilized nn.Linear for efficient processing on my laptop GPU.</p>\n<p>In addition, I introduced a nuanced approach for encoding dense features by breaking them into four distinct encoding layers, each dedicated to capturing nuanced patterns related to mean and standard deviation for both 'sm_name' and 'cell_type'.</p>\n<p>Furthermore, the model architecture leveraged the state-of-the-art Lion optimizer, recognized for its efficacy in optimizing transformer-based models. The specific model implementation, as exemplified in the CustomTransformer_v3 class, showcases a transformer architecture with multiple layers, heads, and dropout for optimal learning.</p>\n<p>In summary, this modularized and intricately designed model not only optimizes computational efficiency but also demonstrates a keen understanding of the unique characteristics of sparse and dense features, contributing to the overall effectiveness of the learning process.</p>\n<p><strong>Hyperparameters</strong><br>\nThe model training regimen was characterized by a judicious selection of hyperparameters, featuring an initial learning rate of 1e-5 and a weight decay of 1e-4, the latter serving as an additional regularization mechanism, particularly for dropout layers within the model architecture. This meticulous choice of hyperparameters aimed at striking a balance between effective training and prevention of overfitting.</p>\n<p>To dynamically adjust the learning rate during training, a learning rate scheduler of the ReduceLROnPlateau type was employed. The scheduler, configured with a mode of \"min\" and a reduction factor of 0.9999, adeptly adapted the learning rate in response to plateaus in model performance, ensuring an efficient convergence towards optimal results. The patience parameter, set to 500, determined the number of epochs with no improvement before triggering a reduction in the learning rate.</p>\n<p>The training process spanned 20,000 epochs, incorporating an early stopping mechanism triggered after 5,000 epochs of stagnation. This comprehensive training strategy underscored a thoughtful balance between fine-tuning the model's parameters and preventing overfitting, contributing to the robustness and efficiency of the learning process.</p>\n<p><strong>Loss function</strong></p>\n<p>While Mean Absolute Error (MAE) and Mean Squared Error (MSE) individually prove effective for gauging bias or variance, it's crucial to recognize their distinct roles in assessing model performance. To strike a balance between these metrics, I opted for Huber loss, amalgamating the strengths of both MAE and MSE. Despite utilizing the Mrrmse metric, the ultimate model selection hinged on the performance demonstrated by the optimal loss function on the validation set.</p>\n<h2>Preventing overfitting</h2>\n<p>In addressing the challenges posed by the limited size of the dataset, a critical consideration was ensuring the model's ability to effectively regularize. To achieve this, I incorporated Dropout layers within the transformer architecture and applied weight decay for L2 penalty on the model weights. Additionally, to mitigate the impact of exploding gradients, I implemented gradient norm clipping with a maximum norm of 1.</p>\n<h2>Validation Strategy</h2>\n<p>In pursuit of an optimal model configuration, I adopted a robust validation strategy by training with different seeds and employing k-fold cross-validation. The diverse set of performance metrics for each fold included:</p>\n<ul>\n<li>Achieving a score of 0.551 through the incorporation of standard deviation, mean, and clustering sampling.</li>\n<li>Attaining a score of 0.559 by excluding uncommon elements from the dataset.</li>\n<li>Realizing a score of 0.575 through clustering sampling.</li>\n<li>Securing a score of 0.554 by integrating mean, random sampling, and excluding standard deviation from the features.<br>\nThis validation approach not only facilitated the identification of the best-performing model but also guided the fine-tuning of hyperparameters to optimize overall model performance.</li>\n</ul>\n<h2>Sources</h2>\n<ul>\n<li><a href=\"https://github.com/lucidrains/lion-pytorch\" target=\"_blank\">https://github.com/lucidrains/lion-pytorch</a></li>\n<li><a href=\"https://www.kaggle.com/code/ayushs9020/understanding-the-competition-open-problems\" target=\"_blank\">https://www.kaggle.com/code/ayushs9020/understanding-the-competition-open-problems</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s</a></li>\n<li><a href=\"https://github.com/Eliorkalfon/single_cell_pb\" target=\"_blank\">https://github.com/Eliorkalfon/single_cell_pb</a></li>\n</ul>",
  "messages": [
    {
      "id": 2545091,
      "postDate": "2023-12-01T09:34:36.337Z",
      "content": "<p>I am thrilled to finally unveil my solution to this competition!</p>\n<h2>Context</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></li>\n</ul>\n<h2>Overview of the approach</h2>\n<p>My optimal model emerged as a composite of four models, where three were employed as a normalized sum, and the fourth acted as an amplifier, showcasing discernible impact, potentially on labels with alternating signs:</p>\n<ul>\n<li><p>weight_df1: 0.5 (utilizing std, mean, and clustering sampling, yielding 0.551)</p></li>\n<li><p>weight_df2: 0.25 (excluding uncommon elements, resulting in 0.559)</p></li>\n<li><p>weight_df3: 0.25 (leveraging clustering sampling, achieving 0.575)</p></li>\n<li><p>weight_df4: 0.3 (incorporating mean, random sampling, and excluding std, attaining 0.554)<br>\nThe resultant model is expressed as the weighted sum of these components: resulted_model = weight_df1 * df1 + weight_df2 * df2 + weight_df3 * df3 + weight_df4 * df4. All models were trained using the same transformer architecture, albeit with varying dimensional models (with the optimal dimension being 128).</p></li>\n</ul>\n<p><strong>Data preprocessing and feature selection</strong><br>\nTo train deep learning models on categorical labels, I employed one-hot encoding transformation. To address high and low bias labels, I utilized target encoding by calculating the mean and standard deviation for each cell type and SM name. During experimentation, I compared models with and without uncommon columns and discovered that including uncommon columns significantly improved the training of encoding layers for mean and standard deviation feature vectors, resulting in noticeable performance enhancement.<br>\n<strong>EDA</strong></p>\n<p>In my feature exploration, I performed truncated Singular Value Decomposition (SVD) specifically on the target variables. Despite experimenting with different SVD sizes, the models generated from this approach consistently fell short of replicating the performance observed in the full targets regression. Notably, as part of this analysis, I identified certain targets with significant standard deviation values.</p>\n<p>To leverage this insight, I strategically enriched the feature set by introducing the standard deviation as a dedicated feature for each cell type and SM name. This involved concatenating these standard deviation values into a unified vector format (std_cell_type, std_sm_name).</p>\n<p>Furthermore, as I delved into the dataset, I conducted a comprehensive examination, revealing a remarkable cleanliness characterized by the absence of NaN values and duplicates. This meticulous exploration not only extended to model training but also encompassed a nuanced analysis of the target variables, contributing to the informed enhancement of the feature set for improved model performance.</p>\n<p><strong>Sampling strategy</strong><br>\nIn devising an effective sampling strategy for partitioning the data into training and validation sets, I employed a sophisticated approach rooted in the clusters derived from a K-Means clustering analysis on the target values. The rationale behind this strategy lies in the premise that data points within similar clusters share inherent patterns or characteristics, thus ensuring a more nuanced and representative split.</p>\n<p>For each distinct cluster identified through K-Means, I executed the data partitioning using the train_test_split function from the scikit-learn library. This method facilitated a controlled allocation of data points to both the training and validation sets, ensuring that the models were exposed to a diverse yet stratified representation of the data.</p>\n<p>The validation percentage, a crucial parameter in this process, was meticulously chosen to strike a balance between model robustness and computational efficiency. By experimenting with validation percentages ranging from 0.1 to 0.2, I systematically assessed the impact on model performance. Ultimately, for the optimal model configuration, a validation percentage of 0.1 was deemed most effective. This decision was guided by a careful consideration of the trade-off between the need for a sufficiently large training set and the importance of a robust validation set to gauge model generalizability.</p>\n<p>In summary, this sampling strategy, grounded in cluster-based data splitting, not only acknowledges the underlying structure within the target values but also optimizes the partitioning process to enhance the model's ability to capture diverse patterns and generalize effectively to unseen data.</p>\n<p><strong>Modeling</strong></p>\n<p>This is my best architecture</p>\n<pre><code>  (nn.Module):  \n     ():\n        (CustomTransformer_v3, self).__init__()\n        self.num_target_encodings =  * \n        self.num_sparse_features = num_features - self.num_target_encodings\n\n        self.sparse_feature_embedding = nn.Linear(self.num_sparse_features, d_model)\n        self.target_encoding_embedding = nn.Linear(self.num_target_encodings, d_model)\n        self.norm = nn.LayerNorm(d_model)\n\n        self.concatenation_layer = nn.Linear( * d_model, d_model)\n        self.transformer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=num_heads, dropout=dropout, activation=nn.GELU(),\n                                       batch_first=),\n            num_layers=num_layers\n        )\n        self.fc = nn.Linear(d_model, num_labels)\n\n     ():\n        sparse_features = x[:, :self.num_sparse_features]\n        target_encodings = x[:, self.num_sparse_features:]\n\n        sparse_features = self.sparse_feature_embedding(sparse_features)\n        target_encodings = self.target_encoding_embedding(target_encodings)\n\n        combined_features = torch.cat((sparse_features, target_encodings), dim=)\n        combined_features = self.concatenation_layer(combined_features)\n        combined_features = self.norm(combined_features)\n\n        x = self.transformer(combined_features)\n        x = self.norm(x)\n\n        x = self.fc(x)\n         x\n</code></pre>\n<p>I modularized my model into two distinct segments to efficiently handle the complexity of the input data:</p>\n<p>Sparse Feature Encoding:</p>\n<p>For encoding sparse features, I utilized an embedding layer tailored to handle the sparsity inherent in certain feature types.<br>\nConsidering the nature of sparse features, I initially explored utilizing the nn.Embedding layer. However, due to computational constraints on my laptop GPU, I opted for an alternative approach using a linear layer (nn.Linear) to convert sparse features into a dense representation.<br>\nTarget Encoding Feature Encoding (Dense Features):</p>\n<p>Concurrently, I addressed the encoding of target encodings, which are inherently dense features.<br>\nTo accomplish this, I employed a separate linear layer (nn.Linear) designed specifically for target encodings, ensuring an effective transformation into a meaningful latent space.<br>\nFollowing these individual encodings, I concatenated the resulting embeddings into a unified latent vector, fostering a comprehensive representation of both sparse and dense feature information. To further enhance the model's capacity to capture nuanced patterns, I employed normalization (nn.LayerNorm) on the concatenated features.</p>\n<p>Considering the computational challenges posed by sparse feature embedding using nn.Embedding, I strategically utilized nn.Linear for efficient processing on my laptop GPU.</p>\n<p>In addition, I introduced a nuanced approach for encoding dense features by breaking them into four distinct encoding layers, each dedicated to capturing nuanced patterns related to mean and standard deviation for both 'sm_name' and 'cell_type'.</p>\n<p>Furthermore, the model architecture leveraged the state-of-the-art Lion optimizer, recognized for its efficacy in optimizing transformer-based models. The specific model implementation, as exemplified in the CustomTransformer_v3 class, showcases a transformer architecture with multiple layers, heads, and dropout for optimal learning.</p>\n<p>In summary, this modularized and intricately designed model not only optimizes computational efficiency but also demonstrates a keen understanding of the unique characteristics of sparse and dense features, contributing to the overall effectiveness of the learning process.</p>\n<p><strong>Hyperparameters</strong><br>\nThe model training regimen was characterized by a judicious selection of hyperparameters, featuring an initial learning rate of 1e-5 and a weight decay of 1e-4, the latter serving as an additional regularization mechanism, particularly for dropout layers within the model architecture. This meticulous choice of hyperparameters aimed at striking a balance between effective training and prevention of overfitting.</p>\n<p>To dynamically adjust the learning rate during training, a learning rate scheduler of the ReduceLROnPlateau type was employed. The scheduler, configured with a mode of \"min\" and a reduction factor of 0.9999, adeptly adapted the learning rate in response to plateaus in model performance, ensuring an efficient convergence towards optimal results. The patience parameter, set to 500, determined the number of epochs with no improvement before triggering a reduction in the learning rate.</p>\n<p>The training process spanned 20,000 epochs, incorporating an early stopping mechanism triggered after 5,000 epochs of stagnation. This comprehensive training strategy underscored a thoughtful balance between fine-tuning the model's parameters and preventing overfitting, contributing to the robustness and efficiency of the learning process.</p>\n<p><strong>Loss function</strong></p>\n<p>While Mean Absolute Error (MAE) and Mean Squared Error (MSE) individually prove effective for gauging bias or variance, it's crucial to recognize their distinct roles in assessing model performance. To strike a balance between these metrics, I opted for Huber loss, amalgamating the strengths of both MAE and MSE. Despite utilizing the Mrrmse metric, the ultimate model selection hinged on the performance demonstrated by the optimal loss function on the validation set.</p>\n<h2>Preventing overfitting</h2>\n<p>In addressing the challenges posed by the limited size of the dataset, a critical consideration was ensuring the model's ability to effectively regularize. To achieve this, I incorporated Dropout layers within the transformer architecture and applied weight decay for L2 penalty on the model weights. Additionally, to mitigate the impact of exploding gradients, I implemented gradient norm clipping with a maximum norm of 1.</p>\n<h2>Validation Strategy</h2>\n<p>In pursuit of an optimal model configuration, I adopted a robust validation strategy by training with different seeds and employing k-fold cross-validation. The diverse set of performance metrics for each fold included:</p>\n<ul>\n<li>Achieving a score of 0.551 through the incorporation of standard deviation, mean, and clustering sampling.</li>\n<li>Attaining a score of 0.559 by excluding uncommon elements from the dataset.</li>\n<li>Realizing a score of 0.575 through clustering sampling.</li>\n<li>Securing a score of 0.554 by integrating mean, random sampling, and excluding standard deviation from the features.<br>\nThis validation approach not only facilitated the identification of the best-performing model but also guided the fine-tuning of hyperparameters to optimize overall model performance.</li>\n</ul>\n<h2>Sources</h2>\n<ul>\n<li><a href=\"https://github.com/lucidrains/lion-pytorch\" target=\"_blank\">https://github.com/lucidrains/lion-pytorch</a></li>\n<li><a href=\"https://www.kaggle.com/code/ayushs9020/understanding-the-competition-open-problems\" target=\"_blank\">https://www.kaggle.com/code/ayushs9020/understanding-the-competition-open-problems</a></li>\n<li><a href=\"https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s</a></li>\n<li><a href=\"https://github.com/Eliorkalfon/single_cell_pb\" target=\"_blank\">https://github.com/Eliorkalfon/single_cell_pb</a></li>\n</ul>",
      "rawMarkdown": "I am thrilled to finally unveil my solution to this competition!\n## Context\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n## Overview of the approach\nMy optimal model emerged as a composite of four models, where three were employed as a normalized sum, and the fourth acted as an amplifier, showcasing discernible impact, potentially on labels with alternating signs:\n\n- weight_df1: 0.5 (utilizing std, mean, and clustering sampling, yielding 0.551)\n- weight_df2: 0.25 (excluding uncommon elements, resulting in 0.559)\n- weight_df3: 0.25 (leveraging clustering sampling, achieving 0.575)\n\n- weight_df4: 0.3 (incorporating mean, random sampling, and excluding std, attaining 0.554)\nThe resultant model is expressed as the weighted sum of these components: resulted_model = weight_df1 * df1 + weight_df2 * df2 + weight_df3 * df3 + weight_df4 * df4. All models were trained using the same transformer architecture, albeit with varying dimensional models (with the optimal dimension being 128).\n\n**Data preprocessing and feature selection**\nTo train deep learning models on categorical labels, I employed one-hot encoding transformation. To address high and low bias labels, I utilized target encoding by calculating the mean and standard deviation for each cell type and SM name. During experimentation, I compared models with and without uncommon columns and discovered that including uncommon columns significantly improved the training of encoding layers for mean and standard deviation feature vectors, resulting in noticeable performance enhancement.\n**EDA**\n\nIn my feature exploration, I performed truncated Singular Value Decomposition (SVD) specifically on the target variables. Despite experimenting with different SVD sizes, the models generated from this approach consistently fell short of replicating the performance observed in the full targets regression. Notably, as part of this analysis, I identified certain targets with significant standard deviation values.\n\nTo leverage this insight, I strategically enriched the feature set by introducing the standard deviation as a dedicated feature for each cell type and SM name. This involved concatenating these standard deviation values into a unified vector format (std_cell_type, std_sm_name).\n\nFurthermore, as I delved into the dataset, I conducted a comprehensive examination, revealing a remarkable cleanliness characterized by the absence of NaN values and duplicates. This meticulous exploration not only extended to model training but also encompassed a nuanced analysis of the target variables, contributing to the informed enhancement of the feature set for improved model performance.\n\n\n**Sampling strategy**\nIn devising an effective sampling strategy for partitioning the data into training and validation sets, I employed a sophisticated approach rooted in the clusters derived from a K-Means clustering analysis on the target values. The rationale behind this strategy lies in the premise that data points within similar clusters share inherent patterns or characteristics, thus ensuring a more nuanced and representative split.\n\nFor each distinct cluster identified through K-Means, I executed the data partitioning using the train_test_split function from the scikit-learn library. This method facilitated a controlled allocation of data points to both the training and validation sets, ensuring that the models were exposed to a diverse yet stratified representation of the data.\n\nThe validation percentage, a crucial parameter in this process, was meticulously chosen to strike a balance between model robustness and computational efficiency. By experimenting with validation percentages ranging from 0.1 to 0.2, I systematically assessed the impact on model performance. Ultimately, for the optimal model configuration, a validation percentage of 0.1 was deemed most effective. This decision was guided by a careful consideration of the trade-off between the need for a sufficiently large training set and the importance of a robust validation set to gauge model generalizability.\n\nIn summary, this sampling strategy, grounded in cluster-based data splitting, not only acknowledges the underlying structure within the target values but also optimizes the partitioning process to enhance the model's ability to capture diverse patterns and generalize effectively to unseen data.\n\n**Modeling**\n\nThis is my best architecture\n```python\n class CustomTransformer_v3(nn.Module):  # mean + std\n    def __init__(self, num_features, num_labels, d_model=128, num_heads=8, num_layers=6, dropout=0.3):\n        super(CustomTransformer_v3, self).__init__()\n        self.num_target_encodings = 18211 * 4\n        self.num_sparse_features = num_features - self.num_target_encodings\n\n        self.sparse_feature_embedding = nn.Linear(self.num_sparse_features, d_model)\n        self.target_encoding_embedding = nn.Linear(self.num_target_encodings, d_model)\n        self.norm = nn.LayerNorm(d_model)\n\n        self.concatenation_layer = nn.Linear(2 * d_model, d_model)\n        self.transformer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=num_heads, dropout=dropout, activation=nn.GELU(),\n                                       batch_first=True),\n            num_layers=num_layers\n        )\n        self.fc = nn.Linear(d_model, num_labels)\n\n    def forward(self, x):\n        sparse_features = x[:, :self.num_sparse_features]\n        target_encodings = x[:, self.num_sparse_features:]\n\n        sparse_features = self.sparse_feature_embedding(sparse_features)\n        target_encodings = self.target_encoding_embedding(target_encodings)\n\n        combined_features = torch.cat((sparse_features, target_encodings), dim=1)\n        combined_features = self.concatenation_layer(combined_features)\n        combined_features = self.norm(combined_features)\n\n        x = self.transformer(combined_features)\n        x = self.norm(x)\n\n        x = self.fc(x)\n        return x\n\n```\nI modularized my model into two distinct segments to efficiently handle the complexity of the input data:\n\nSparse Feature Encoding:\n\nFor encoding sparse features, I utilized an embedding layer tailored to handle the sparsity inherent in certain feature types.\nConsidering the nature of sparse features, I initially explored utilizing the nn.Embedding layer. However, due to computational constraints on my laptop GPU, I opted for an alternative approach using a linear layer (nn.Linear) to convert sparse features into a dense representation.\nTarget Encoding Feature Encoding (Dense Features):\n\nConcurrently, I addressed the encoding of target encodings, which are inherently dense features.\nTo accomplish this, I employed a separate linear layer (nn.Linear) designed specifically for target encodings, ensuring an effective transformation into a meaningful latent space.\nFollowing these individual encodings, I concatenated the resulting embeddings into a unified latent vector, fostering a comprehensive representation of both sparse and dense feature information. To further enhance the model's capacity to capture nuanced patterns, I employed normalization (nn.LayerNorm) on the concatenated features.\n\nConsidering the computational challenges posed by sparse feature embedding using nn.Embedding, I strategically utilized nn.Linear for efficient processing on my laptop GPU.\n\nIn addition, I introduced a nuanced approach for encoding dense features by breaking them into four distinct encoding layers, each dedicated to capturing nuanced patterns related to mean and standard deviation for both 'sm_name' and 'cell_type'.\n\nFurthermore, the model architecture leveraged the state-of-the-art Lion optimizer, recognized for its efficacy in optimizing transformer-based models. The specific model implementation, as exemplified in the CustomTransformer_v3 class, showcases a transformer architecture with multiple layers, heads, and dropout for optimal learning.\n\nIn summary, this modularized and intricately designed model not only optimizes computational efficiency but also demonstrates a keen understanding of the unique characteristics of sparse and dense features, contributing to the overall effectiveness of the learning process.\n\n\n**Hyperparameters**\nThe model training regimen was characterized by a judicious selection of hyperparameters, featuring an initial learning rate of 1e-5 and a weight decay of 1e-4, the latter serving as an additional regularization mechanism, particularly for dropout layers within the model architecture. This meticulous choice of hyperparameters aimed at striking a balance between effective training and prevention of overfitting.\n\nTo dynamically adjust the learning rate during training, a learning rate scheduler of the ReduceLROnPlateau type was employed. The scheduler, configured with a mode of \"min\" and a reduction factor of 0.9999, adeptly adapted the learning rate in response to plateaus in model performance, ensuring an efficient convergence towards optimal results. The patience parameter, set to 500, determined the number of epochs with no improvement before triggering a reduction in the learning rate.\n\nThe training process spanned 20,000 epochs, incorporating an early stopping mechanism triggered after 5,000 epochs of stagnation. This comprehensive training strategy underscored a thoughtful balance between fine-tuning the model's parameters and preventing overfitting, contributing to the robustness and efficiency of the learning process.\n\n**Loss function**\n\nWhile Mean Absolute Error (MAE) and Mean Squared Error (MSE) individually prove effective for gauging bias or variance, it's crucial to recognize their distinct roles in assessing model performance. To strike a balance between these metrics, I opted for Huber loss, amalgamating the strengths of both MAE and MSE. Despite utilizing the Mrrmse metric, the ultimate model selection hinged on the performance demonstrated by the optimal loss function on the validation set.\n\n## Preventing overfitting\nIn addressing the challenges posed by the limited size of the dataset, a critical consideration was ensuring the model's ability to effectively regularize. To achieve this, I incorporated Dropout layers within the transformer architecture and applied weight decay for L2 penalty on the model weights. Additionally, to mitigate the impact of exploding gradients, I implemented gradient norm clipping with a maximum norm of 1.\n\n## Validation Strategy\nIn pursuit of an optimal model configuration, I adopted a robust validation strategy by training with different seeds and employing k-fold cross-validation. The diverse set of performance metrics for each fold included:\n\n- Achieving a score of 0.551 through the incorporation of standard deviation, mean, and clustering sampling.\n- Attaining a score of 0.559 by excluding uncommon elements from the dataset.\n- Realizing a score of 0.575 through clustering sampling.\n- Securing a score of 0.554 by integrating mean, random sampling, and excluding standard deviation from the features.\nThis validation approach not only facilitated the identification of the best-performing model but also guided the fine-tuning of hyperparameters to optimize overall model performance.\n## Sources\n- https://github.com/lucidrains/lion-pytorch\n- https://www.kaggle.com/code/ayushs9020/understanding-the-competition-open-problems\n- https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\n- https://github.com/Eliorkalfon/single_cell_pb\n",
      "votes": 49
    },
    {
      "id": 2545964,
      "postDate": "2023-12-02T01:32:45.897Z",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> 🎉🎉🎉Thank you very much for sharing. Your work is so creative. But due to my limited knowledge, I still cannot understand some parts well. I am already following your code sharing. Looking forward to your future achievements!</p>",
      "rawMarkdown": "Congratulations! @eliork 🎉🎉🎉Thank you very much for sharing. Your work is so creative. But due to my limited knowledge, I still cannot understand some parts well. I am already following your code sharing. Looking forward to your future achievements!",
      "votes": 6,
      "replies": [
        {
          "id": 2547275,
          "postDate": "2023-12-03T11:16:38.170Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/songqizhou\" target=\"_blank\">@songqizhou</a>, I've uploaded the code to the Git repository. Please feel free to explore it and don't hesitate to reach out if you have any questions.</p>",
          "rawMarkdown": "Hi @songqizhou, I've uploaded the code to the Git repository. Please feel free to explore it and don't hesitate to reach out if you have any questions.",
          "votes": 1,
          "replies": [
            {
              "id": 2548184,
              "postDate": "2023-12-04T07:37:23.963Z",
              "content": "<p>Many thanks to you <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a>! Your code is as elegant as poetry!🥰</p>",
              "rawMarkdown": "Many thanks to you @eliork! Your code is as elegant as poetry!🥰",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2545111,
      "postDate": "2023-12-01T09:55:10.963Z",
      "content": "<p>congrats !! <br>\ni thought none of the top solution would be transformer based</p>",
      "rawMarkdown": "congrats !! \ni thought none of the top solution would be transformer based\n",
      "votes": 3,
      "replies": [
        {
          "id": 2545172,
          "postDate": "2023-12-01T10:46:56.813Z",
          "content": "<p>Thanks! I tried xgboost and simple mlp  but the results were the best for this model. </p>",
          "rawMarkdown": "Thanks! I tried xgboost and simple mlp  but the results were the best for this model. ",
          "votes": 2
        },
        {
          "id": 2546938,
          "postDate": "2023-12-03T02:41:45.023Z",
          "content": "<p>This competition made me realize the power of the \"generative\" NN model, generating 18211 points from chem with 155 ranks and cell with 6 ranks, though the smart feature engineering makes the difference.</p>",
          "rawMarkdown": "This competition made me realize the power of the \"generative\" NN model, generating 18211 points from chem with 155 ranks and cell with 6 ranks, though the smart feature engineering makes the difference.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2556380,
      "postDate": "2023-12-10T16:07:52.213Z",
      "content": "<p>Congratulations with the great result ! <br>\nWould you be so kind to share the submission files of your solutions - for further research analysis  ? </p>",
      "rawMarkdown": "Congratulations with the great result ! \nWould you be so kind to share the submission files of your solutions - for further research analysis  ? \n",
      "votes": 1,
      "replies": [
        {
          "id": 2556388,
          "postDate": "2023-12-10T16:15:07.673Z",
          "content": "<p>Thanks! The files are in my GitHub repository under submissions folder also check seq.py for more info. </p>",
          "rawMarkdown": "Thanks! The files are in my GitHub repository under submissions folder also check seq.py for more info. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2549515,
      "postDate": "2023-12-05T10:05:19.620Z",
      "content": "<p>I followed you during the competition and I was sure you would be recognized in the private LB. Superb job <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> ! :)</p>",
      "rawMarkdown": "I followed you during the competition and I was sure you would be recognized in the private LB. Superb job @eliork ! :)",
      "votes": 1
    },
    {
      "id": 2547767,
      "postDate": "2023-12-03T20:05:00.767Z",
      "content": "<p>Thank you a lot for sharing your approach! And congratulations on the excellent result!</p>",
      "rawMarkdown": "Thank you a lot for sharing your approach! And congratulations on the excellent result!",
      "votes": 1
    },
    {
      "id": 2547308,
      "postDate": "2023-12-03T12:08:35.827Z",
      "content": "<p>Congratulations on achieving the 2nd place. Thanks for sharing details about CustomTransformer model.  Very interesting indeed.</p>",
      "rawMarkdown": "Congratulations on achieving the 2nd place. Thanks for sharing details about CustomTransformer model.  Very interesting indeed.",
      "votes": 1
    },
    {
      "id": 2546311,
      "postDate": "2023-12-02T09:42:50.793Z",
      "content": "<p>Great Efforts Really!</p>",
      "rawMarkdown": "Great Efforts Really!",
      "votes": 1
    },
    {
      "id": 2545286,
      "postDate": "2023-12-01T12:08:27.960Z",
      "content": "<p>Congratulations and thank you very much for sharing the solution!  Seeing your solution is indeed an aha moment; it is unbiased and summarizing both cell and chem features; wow, simple and strong in prediction, it is quite a learning.</p>",
      "rawMarkdown": "Congratulations and thank you very much for sharing the solution!  Seeing your solution is indeed an aha moment; it is unbiased and summarizing both cell and chem features; wow, simple and strong in prediction, it is quite a learning.",
      "votes": 2
    },
    {
      "id": 2545173,
      "postDate": "2023-12-01T10:48:32.767Z",
      "content": "<p>I have been dreaming of seeing your solutions for a long time.</p>",
      "rawMarkdown": "I have been dreaming of seeing your solutions for a long time.",
      "votes": 2,
      "replies": [
        {
          "id": 2545200,
          "postDate": "2023-12-01T11:00:59.393Z",
          "content": "<p>I hope I didn’t disappoint  :)</p>",
          "rawMarkdown": "I hope I didn’t disappoint  :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 2545580,
      "postDate": "2023-12-01T15:24:15.043Z",
      "content": "<p>Your notebooks are truly commendable! I upvoted all of them. If you get a chance, I'd be grateful for your input on my latest work. Thanks</p>",
      "rawMarkdown": "Your notebooks are truly commendable! I upvoted all of them. If you get a chance, I'd be grateful for your input on my latest work. Thanks"
    },
    {
      "id": 2608084,
      "postDate": "2024-01-18T14:56:00.903Z",
      "content": "<p>congrats !!</p>\n<p>I‘m confused about the discription : [My optimal model emerged as a composite of four models], aren't these models trained seperately? </p>\n<p>How can we conbine them together?  Or just compute the weighted sum of their separate results?</p>\n<p>Thanks a lot for your contributions!</p>",
      "rawMarkdown": "congrats !!\n\nI‘m confused about the discription : [My optimal model emerged as a composite of four models], aren't these models trained seperately? \n\nHow can we conbine them together?  Or just compute the weighted sum of their separate results?\n\nThanks a lot for your contributions!",
      "replies": [
        {
          "id": 2608192,
          "postDate": "2024-01-18T15:46:33.143Z",
          "content": "<p>The models were trained separately, with the same general architecture but with different features (mean,std,etc), I just ensembled the results using weighted sum, I hope it clarifies my approach. </p>",
          "rawMarkdown": "The models were trained separately, with the same general architecture but with different features (mean,std,etc), I just ensembled the results using weighted sum, I hope it clarifies my approach. ",
          "replies": [
            {
              "id": 2608810,
              "postDate": "2024-01-19T05:41:58.703Z",
              "content": "<p>Thanks for your reply!</p>\n<p>I've checked the seq.py，but I still have some questions about the specific features of weight_df.</p>\n<p>My understands are as follows, please correct me if I'm wrong:</p>\n<ul>\n<li><p><strong>weight_df1</strong>: Utilizing std, mean, and clustering sampling</p></li>\n<li><p><strong>weight_df2</strong>:  Same as weight_df1 but excluding uncommon elements.</p></li>\n<li><p><strong>weight_df3</strong>:  I'm not quite understand how to ' <strong>leveraging</strong> ' clustering sampling, are you using std and mean both here? <br>\nWhat are the differences between Weight_ df1 and weight_df3 ?</p></li>\n<li><p><strong>weight_df4</strong>: Incorporating mean, random sampling, and excluding std</p></li>\n</ul>\n<p>Thanks again for your reply!</p>",
              "rawMarkdown": "Thanks for your reply!\n\nI've checked the seq.py，but I still have some questions about the specific features of weight_df.\n\nMy understands are as follows, please correct me if I'm wrong:\n\n- **weight_df1**: Utilizing std, mean, and clustering sampling\n\n- **weight_df2**:  Same as weight_df1 but excluding uncommon elements.\n\n- **weight_df3**:  I'm not quite understand how to ' **leveraging** ' clustering sampling, are you using std and mean both here? \nWhat are the differences between Weight_ df1 and weight_df3 ?\n\n- **weight_df4**: Incorporating mean, random sampling, and excluding std\n\nThanks again for your reply!"
            },
            {
              "id": 2609034,
              "postDate": "2024-01-19T08:10:14.097Z",
              "content": "<p>The difference between df1 to df3 is the sampling method.<br>\nweight_df3: I clustered all of the data using k-means, and I sampled from each cluster instead of random sampling.</p>",
              "rawMarkdown": "The difference between df1 to df3 is the sampling method.\nweight_df3: I clustered all of the data using k-means, and I sampled from each cluster instead of random sampling."
            }
          ]
        }
      ]
    },
    {
      "id": 2549326,
      "postDate": "2023-12-05T06:24:11.627Z",
      "content": "<p>This is my brief summary on the work flow of the solution (correct me if I'm wrong🤣):</p>\n<ul>\n<li>data augmentation: utilized std and mean of the <code>train_data</code> and turned the train_data shape from <code>(614,152)</code> to <code>(614,72996)</code>, y shape to <code>(614,18211)</code>.</li>\n<li>then utilize K-means to clean the train data?</li>\n<li>then (check the model) train the transformer model (with 2 sperated MLP for sparse and dense features then a concatenate) with Lion and HuberLoss, output a rather high dim predictions.</li>\n<li>predict the test data using truancatedSVD to down sample the predictions to submission format.</li>\n</ul>",
      "rawMarkdown": "This is my brief summary on the work flow of the solution (correct me if I'm wrong🤣):\n- data augmentation: utilized std and mean of the ``train_data`` and turned the train_data shape from ``(614,152)`` to ``(614,72996)``, y shape to ``(614,18211)``.\n- then utilize K-means to clean the train data?\n- then (check the model) train the transformer model (with 2 sperated MLP for sparse and dense features then a concatenate) with Lion and HuberLoss, output a rather high dim predictions.\n- predict the test data using truancatedSVD to down sample the predictions to submission format.\n\n",
      "replies": [
        {
          "id": 2549345,
          "postDate": "2023-12-05T06:50:59.850Z",
          "content": "<p>I employed K-means to create samples from each cluster. A simple analogy is when dealing with significantly imbalanced data, it's crucial to ensure a representation of categories in both the training and validation sets. The optimal model was attained without truncated SVD. However, if you choose to apply this transformation, remember to also perform the inverse transform.</p>",
          "rawMarkdown": "I employed K-means to create samples from each cluster. A simple analogy is when dealing with significantly imbalanced data, it's crucial to ensure a representation of categories in both the training and validation sets. The optimal model was attained without truncated SVD. However, if you choose to apply this transformation, remember to also perform the inverse transform.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2546667,
      "postDate": "2023-12-02T17:19:51.950Z",
      "content": "<p>Can you give me an example of when you utilized the Lion-Python optimizer?</p>\n<p>I have been waiting for the model.</p>\n<blockquote>\n  <p>Deep learning models for the Kaggle's Open Problems – Single-Cell Perturbations competition. Coming soon. <br>\n  <a href=\"https://github.com/Eliorkalfon/single_cell_pb\" target=\"_blank\">https://github.com/Eliorkalfon/single_cell_pb</a>)</p>\n</blockquote>",
      "rawMarkdown": "Can you give me an example of when you utilized the Lion-Python optimizer?\n\nI have been waiting for the model.\n>Deep learning models for the Kaggle's Open Problems – Single-Cell Perturbations competition. Coming soon. \n>https://github.com/Eliorkalfon/single_cell_pb)",
      "replies": [
        {
          "id": 2547278,
          "postDate": "2023-12-03T11:17:31.653Z",
          "content": "<p>Hi, I've uploaded the code to the Git repository, check it out :)</p>",
          "rawMarkdown": "Hi, I've uploaded the code to the Git repository, check it out :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2559121,
      "postDate": "2023-12-12T16:07:24.337Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2548329,
      "postDate": "2023-12-04T10:33:06.630Z",
      "content": "<p>Congrats! Thanks for sharing </p>",
      "rawMarkdown": "Congrats! Thanks for sharing ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2545964,
      "author_name": "BarryZhou",
      "author_url": "",
      "post_date": "2023-12-02T01:32:45.897000",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> 🎉🎉🎉Thank you very much for sharing. Your work is so creative. But due to my limited knowledge, I still cannot understand some parts well. I am already following your code sharing. Looking forward to your future achievements!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2547275,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2023-12-03T11:16:38.170000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/songqizhou\" target=\"_blank\">@songqizhou</a>, I've uploaded the code to the Git repository. Please feel free to explore it and don't hesitate to reach out if you have any questions.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2548184,
              "author_name": "BarryZhou",
              "author_url": "",
              "post_date": "2023-12-04T07:37:23.963000",
              "content": "<p>Many thanks to you <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a>! Your code is as elegant as poetry!🥰</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2545111,
      "author_name": "Séraphin Lampion",
      "author_url": "",
      "post_date": "2023-12-01T09:55:10.963000",
      "content": "<p>congrats !! <br>\ni thought none of the top solution would be transformer based</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2545172,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2023-12-01T10:46:56.813000",
          "content": "<p>Thanks! I tried xgboost and simple mlp  but the results were the best for this model. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2546938,
          "author_name": "makio323",
          "author_url": "",
          "post_date": "2023-12-03T02:41:45.023000",
          "content": "<p>This competition made me realize the power of the \"generative\" NN model, generating 18211 points from chem with 155 ranks and cell with 6 ranks, though the smart feature engineering makes the difference.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2556380,
      "author_name": "Alexander Chervov",
      "author_url": "",
      "post_date": "2023-12-10T16:07:52.213000",
      "content": "<p>Congratulations with the great result ! <br>\nWould you be so kind to share the submission files of your solutions - for further research analysis  ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2556388,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2023-12-10T16:15:07.673000",
          "content": "<p>Thanks! The files are in my GitHub repository under submissions folder also check seq.py for more info. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2549515,
      "author_name": "Octavi Grau",
      "author_url": "",
      "post_date": "2023-12-05T10:05:19.620000",
      "content": "<p>I followed you during the competition and I was sure you would be recognized in the private LB. Superb job <a href=\"https://www.kaggle.com/eliork\" target=\"_blank\">@eliork</a> ! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2547767,
      "author_name": "Aya",
      "author_url": "",
      "post_date": "2023-12-03T20:05:00.767000",
      "content": "<p>Thank you a lot for sharing your approach! And congratulations on the excellent result!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2547308,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-12-03T12:08:35.827000",
      "content": "<p>Congratulations on achieving the 2nd place. Thanks for sharing details about CustomTransformer model.  Very interesting indeed.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2546311,
      "author_name": "Abhas Malguri",
      "author_url": "",
      "post_date": "2023-12-02T09:42:50.793000",
      "content": "<p>Great Efforts Really!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2545286,
      "author_name": "makio323",
      "author_url": "",
      "post_date": "2023-12-01T12:08:27.960000",
      "content": "<p>Congratulations and thank you very much for sharing the solution!  Seeing your solution is indeed an aha moment; it is unbiased and summarizing both cell and chem features; wow, simple and strong in prediction, it is quite a learning.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2545173,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-01T10:48:32.767000",
      "content": "<p>I have been dreaming of seeing your solutions for a long time.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2545200,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2023-12-01T11:00:59.393000",
          "content": "<p>I hope I didn’t disappoint  :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2545580,
      "author_name": "VIKRAM MISHRA",
      "author_url": "",
      "post_date": "2023-12-01T15:24:15.043000",
      "content": "<p>Your notebooks are truly commendable! I upvoted all of them. If you get a chance, I'd be grateful for your input on my latest work. Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2608084,
      "author_name": "yu3jun",
      "author_url": "",
      "post_date": "2024-01-18T14:56:00.903000",
      "content": "<p>congrats !!</p>\n<p>I‘m confused about the discription : [My optimal model emerged as a composite of four models], aren't these models trained seperately? </p>\n<p>How can we conbine them together?  Or just compute the weighted sum of their separate results?</p>\n<p>Thanks a lot for your contributions!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2608192,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2024-01-18T15:46:33.143000",
          "content": "<p>The models were trained separately, with the same general architecture but with different features (mean,std,etc), I just ensembled the results using weighted sum, I hope it clarifies my approach. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2608810,
              "author_name": "yu3jun",
              "author_url": "",
              "post_date": "2024-01-19T05:41:58.703000",
              "content": "<p>Thanks for your reply!</p>\n<p>I've checked the seq.py，but I still have some questions about the specific features of weight_df.</p>\n<p>My understands are as follows, please correct me if I'm wrong:</p>\n<ul>\n<li><p><strong>weight_df1</strong>: Utilizing std, mean, and clustering sampling</p></li>\n<li><p><strong>weight_df2</strong>:  Same as weight_df1 but excluding uncommon elements.</p></li>\n<li><p><strong>weight_df3</strong>:  I'm not quite understand how to ' <strong>leveraging</strong> ' clustering sampling, are you using std and mean both here? <br>\nWhat are the differences between Weight_ df1 and weight_df3 ?</p></li>\n<li><p><strong>weight_df4</strong>: Incorporating mean, random sampling, and excluding std</p></li>\n</ul>\n<p>Thanks again for your reply!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2609034,
              "author_name": "EliKal",
              "author_url": "",
              "post_date": "2024-01-19T08:10:14.097000",
              "content": "<p>The difference between df1 to df3 is the sampling method.<br>\nweight_df3: I clustered all of the data using k-means, and I sampled from each cluster instead of random sampling.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2549326,
      "author_name": "Yonggie",
      "author_url": "",
      "post_date": "2023-12-05T06:24:11.627000",
      "content": "<p>This is my brief summary on the work flow of the solution (correct me if I'm wrong🤣):</p>\n<ul>\n<li>data augmentation: utilized std and mean of the <code>train_data</code> and turned the train_data shape from <code>(614,152)</code> to <code>(614,72996)</code>, y shape to <code>(614,18211)</code>.</li>\n<li>then utilize K-means to clean the train data?</li>\n<li>then (check the model) train the transformer model (with 2 sperated MLP for sparse and dense features then a concatenate) with Lion and HuberLoss, output a rather high dim predictions.</li>\n<li>predict the test data using truancatedSVD to down sample the predictions to submission format.</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2549345,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2023-12-05T06:50:59.850000",
          "content": "<p>I employed K-means to create samples from each cluster. A simple analogy is when dealing with significantly imbalanced data, it's crucial to ensure a representation of categories in both the training and validation sets. The optimal model was attained without truncated SVD. However, if you choose to apply this transformation, remember to also perform the inverse transform.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2546667,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-02T17:19:51.950000",
      "content": "<p>Can you give me an example of when you utilized the Lion-Python optimizer?</p>\n<p>I have been waiting for the model.</p>\n<blockquote>\n  <p>Deep learning models for the Kaggle's Open Problems – Single-Cell Perturbations competition. Coming soon. <br>\n  <a href=\"https://github.com/Eliorkalfon/single_cell_pb\" target=\"_blank\">https://github.com/Eliorkalfon/single_cell_pb</a>)</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 2547278,
          "author_name": "EliKal",
          "author_url": "",
          "post_date": "2023-12-03T11:17:31.653000",
          "content": "<p>Hi, I've uploaded the code to the Git repository, check it out :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2559121,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T16:07:24.337000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2548329,
      "author_name": "hannah142",
      "author_url": "",
      "post_date": "2023-12-04T10:33:06.630000",
      "content": "<p>Congrats! Thanks for sharing </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2545091": "I am thrilled to finally unveil my solution to this competition!\n## Context\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n## Overview of the approach\nMy optimal model emerged as a composite of four models, where three were employed as a normalized sum, and the fourth acted as an amplifier, showcasing discernible impact, potentially on labels with alternating signs:\n\n- weight_df1: 0.5 (utilizing std, mean, and clustering sampling, yielding 0.551)\n- weight_df2: 0.25 (excluding uncommon elements, resulting in 0.559)\n- weight_df3: 0.25 (leveraging clustering sampling, achieving 0.575)\n\n- weight_df4: 0.3 (incorporating mean, random sampling, and excluding std, attaining 0.554)\nThe resultant model is expressed as the weighted sum of these components: resulted_model = weight_df1 * df1 + weight_df2 * df2 + weight_df3 * df3 + weight_df4 * df4. All models were trained using the same transformer architecture, albeit with varying dimensional models (with the optimal dimension being 128).\n\n**Data preprocessing and feature selection**\nTo train deep learning models on categorical labels, I employed one-hot encoding transformation. To address high and low bias labels, I utilized target encoding by calculating the mean and standard deviation for each cell type and SM name. During experimentation, I compared models with and without uncommon columns and discovered that including uncommon columns significantly improved the training of encoding layers for mean and standard deviation feature vectors, resulting in noticeable performance enhancement.\n**EDA**\n\nIn my feature exploration, I performed truncated Singular Value Decomposition (SVD) specifically on the target variables. Despite experimenting with different SVD sizes, the models generated from this approach consistently fell short of replicating the performance observed in the full targets regression. Notably, as part of this analysis, I identified certain targets with significant standard deviation values.\n\nTo leverage this insight, I strategically enriched the feature set by introducing the standard deviation as a dedicated feature for each cell type and SM name. This involved concatenating these standard deviation values into a unified vector format (std_cell_type, std_sm_name).\n\nFurthermore, as I delved into the dataset, I conducted a comprehensive examination, revealing a remarkable cleanliness characterized by the absence of NaN values and duplicates. This meticulous exploration not only extended to model training but also encompassed a nuanced analysis of the target variables, contributing to the informed enhancement of the feature set for improved model performance.\n\n\n**Sampling strategy**\nIn devising an effective sampling strategy for partitioning the data into training and validation sets, I employed a sophisticated approach rooted in the clusters derived from a K-Means clustering analysis on the target values. The rationale behind this strategy lies in the premise that data points within similar clusters share inherent patterns or characteristics, thus ensuring a more nuanced and representative split.\n\nFor each distinct cluster identified through K-Means, I executed the data partitioning using the train_test_split function from the scikit-learn library. This method facilitated a controlled allocation of data points to both the training and validation sets, ensuring that the models were exposed to a diverse yet stratified representation of the data.\n\nThe validation percentage, a crucial parameter in this process, was meticulously chosen to strike a balance between model robustness and computational efficiency. By experimenting with validation percentages ranging from 0.1 to 0.2, I systematically assessed the impact on model performance. Ultimately, for the optimal model configuration, a validation percentage of 0.1 was deemed most effective. This decision was guided by a careful consideration of the trade-off between the need for a sufficiently large training set and the importance of a robust validation set to gauge model generalizability.\n\nIn summary, this sampling strategy, grounded in cluster-based data splitting, not only acknowledges the underlying structure within the target values but also optimizes the partitioning process to enhance the model's ability to capture diverse patterns and generalize effectively to unseen data.\n\n**Modeling**\n\nThis is my best architecture\n```python\n class CustomTransformer_v3(nn.Module):  # mean + std\n    def __init__(self, num_features, num_labels, d_model=128, num_heads=8, num_layers=6, dropout=0.3):\n        super(CustomTransformer_v3, self).__init__()\n        self.num_target_encodings = 18211 * 4\n        self.num_sparse_features = num_features - self.num_target_encodings\n\n        self.sparse_feature_embedding = nn.Linear(self.num_sparse_features, d_model)\n        self.target_encoding_embedding = nn.Linear(self.num_target_encodings, d_model)\n        self.norm = nn.LayerNorm(d_model)\n\n        self.concatenation_layer = nn.Linear(2 * d_model, d_model)\n        self.transformer = nn.TransformerEncoder(\n            nn.TransformerEncoderLayer(d_model=d_model, nhead=num_heads, dropout=dropout, activation=nn.GELU(),\n                                       batch_first=True),\n            num_layers=num_layers\n        )\n        self.fc = nn.Linear(d_model, num_labels)\n\n    def forward(self, x):\n        sparse_features = x[:, :self.num_sparse_features]\n        target_encodings = x[:, self.num_sparse_features:]\n\n        sparse_features = self.sparse_feature_embedding(sparse_features)\n        target_encodings = self.target_encoding_embedding(target_encodings)\n\n        combined_features = torch.cat((sparse_features, target_encodings), dim=1)\n        combined_features = self.concatenation_layer(combined_features)\n        combined_features = self.norm(combined_features)\n\n        x = self.transformer(combined_features)\n        x = self.norm(x)\n\n        x = self.fc(x)\n        return x\n\n```\nI modularized my model into two distinct segments to efficiently handle the complexity of the input data:\n\nSparse Feature Encoding:\n\nFor encoding sparse features, I utilized an embedding layer tailored to handle the sparsity inherent in certain feature types.\nConsidering the nature of sparse features, I initially explored utilizing the nn.Embedding layer. However, due to computational constraints on my laptop GPU, I opted for an alternative approach using a linear layer (nn.Linear) to convert sparse features into a dense representation.\nTarget Encoding Feature Encoding (Dense Features):\n\nConcurrently, I addressed the encoding of target encodings, which are inherently dense features.\nTo accomplish this, I employed a separate linear layer (nn.Linear) designed specifically for target encodings, ensuring an effective transformation into a meaningful latent space.\nFollowing these individual encodings, I concatenated the resulting embeddings into a unified latent vector, fostering a comprehensive representation of both sparse and dense feature information. To further enhance the model's capacity to capture nuanced patterns, I employed normalization (nn.LayerNorm) on the concatenated features.\n\nConsidering the computational challenges posed by sparse feature embedding using nn.Embedding, I strategically utilized nn.Linear for efficient processing on my laptop GPU.\n\nIn addition, I introduced a nuanced approach for encoding dense features by breaking them into four distinct encoding layers, each dedicated to capturing nuanced patterns related to mean and standard deviation for both 'sm_name' and 'cell_type'.\n\nFurthermore, the model architecture leveraged the state-of-the-art Lion optimizer, recognized for its efficacy in optimizing transformer-based models. The specific model implementation, as exemplified in the CustomTransformer_v3 class, showcases a transformer architecture with multiple layers, heads, and dropout for optimal learning.\n\nIn summary, this modularized and intricately designed model not only optimizes computational efficiency but also demonstrates a keen understanding of the unique characteristics of sparse and dense features, contributing to the overall effectiveness of the learning process.\n\n\n**Hyperparameters**\nThe model training regimen was characterized by a judicious selection of hyperparameters, featuring an initial learning rate of 1e-5 and a weight decay of 1e-4, the latter serving as an additional regularization mechanism, particularly for dropout layers within the model architecture. This meticulous choice of hyperparameters aimed at striking a balance between effective training and prevention of overfitting.\n\nTo dynamically adjust the learning rate during training, a learning rate scheduler of the ReduceLROnPlateau type was employed. The scheduler, configured with a mode of \"min\" and a reduction factor of 0.9999, adeptly adapted the learning rate in response to plateaus in model performance, ensuring an efficient convergence towards optimal results. The patience parameter, set to 500, determined the number of epochs with no improvement before triggering a reduction in the learning rate.\n\nThe training process spanned 20,000 epochs, incorporating an early stopping mechanism triggered after 5,000 epochs of stagnation. This comprehensive training strategy underscored a thoughtful balance between fine-tuning the model's parameters and preventing overfitting, contributing to the robustness and efficiency of the learning process.\n\n**Loss function**\n\nWhile Mean Absolute Error (MAE) and Mean Squared Error (MSE) individually prove effective for gauging bias or variance, it's crucial to recognize their distinct roles in assessing model performance. To strike a balance between these metrics, I opted for Huber loss, amalgamating the strengths of both MAE and MSE. Despite utilizing the Mrrmse metric, the ultimate model selection hinged on the performance demonstrated by the optimal loss function on the validation set.\n\n## Preventing overfitting\nIn addressing the challenges posed by the limited size of the dataset, a critical consideration was ensuring the model's ability to effectively regularize. To achieve this, I incorporated Dropout layers within the transformer architecture and applied weight decay for L2 penalty on the model weights. Additionally, to mitigate the impact of exploding gradients, I implemented gradient norm clipping with a maximum norm of 1.\n\n## Validation Strategy\nIn pursuit of an optimal model configuration, I adopted a robust validation strategy by training with different seeds and employing k-fold cross-validation. The diverse set of performance metrics for each fold included:\n\n- Achieving a score of 0.551 through the incorporation of standard deviation, mean, and clustering sampling.\n- Attaining a score of 0.559 by excluding uncommon elements from the dataset.\n- Realizing a score of 0.575 through clustering sampling.\n- Securing a score of 0.554 by integrating mean, random sampling, and excluding standard deviation from the features.\nThis validation approach not only facilitated the identification of the best-performing model but also guided the fine-tuning of hyperparameters to optimize overall model performance.\n## Sources\n- https://github.com/lucidrains/lion-pytorch\n- https://www.kaggle.com/code/ayushs9020/understanding-the-competition-open-problems\n- https://www.kaggle.com/code/alexandervc/op2-eda-baseline-s\n- https://github.com/Eliorkalfon/single_cell_pb\n",
    "2545964": "Congratulations! @eliork 🎉🎉🎉Thank you very much for sharing. Your work is so creative. But due to my limited knowledge, I still cannot understand some parts well. I am already following your code sharing. Looking forward to your future achievements!",
    "2545111": "congrats !! \ni thought none of the top solution would be transformer based\n",
    "2556380": "Congratulations with the great result ! \nWould you be so kind to share the submission files of your solutions - for further research analysis  ? \n",
    "2549515": "I followed you during the competition and I was sure you would be recognized in the private LB. Superb job @eliork ! :)",
    "2547767": "Thank you a lot for sharing your approach! And congratulations on the excellent result!",
    "2547308": "Congratulations on achieving the 2nd place. Thanks for sharing details about CustomTransformer model.  Very interesting indeed.",
    "2546311": "Great Efforts Really!",
    "2545286": "Congratulations and thank you very much for sharing the solution!  Seeing your solution is indeed an aha moment; it is unbiased and summarizing both cell and chem features; wow, simple and strong in prediction, it is quite a learning.",
    "2545173": "I have been dreaming of seeing your solutions for a long time.",
    "2545580": "Your notebooks are truly commendable! I upvoted all of them. If you get a chance, I'd be grateful for your input on my latest work. Thanks",
    "2608084": "congrats !!\n\nI‘m confused about the discription : [My optimal model emerged as a composite of four models], aren't these models trained seperately? \n\nHow can we conbine them together?  Or just compute the weighted sum of their separate results?\n\nThanks a lot for your contributions!",
    "2549326": "This is my brief summary on the work flow of the solution (correct me if I'm wrong🤣):\n- data augmentation: utilized std and mean of the ``train_data`` and turned the train_data shape from ``(614,152)`` to ``(614,72996)``, y shape to ``(614,18211)``.\n- then utilize K-means to clean the train data?\n- then (check the model) train the transformer model (with 2 sperated MLP for sparse and dense features then a concatenate) with Lion and HuberLoss, output a rather high dim predictions.\n- predict the test data using truancatedSVD to down sample the predictions to submission format.\n\n",
    "2546667": "Can you give me an example of when you utilized the Lion-Python optimizer?\n\nI have been waiting for the model.\n>Deep learning models for the Kaggle's Open Problems – Single-Cell Perturbations competition. Coming soon. \n>https://github.com/Eliorkalfon/single_cell_pb)",
    "2559121": "",
    "2548329": "Congrats! Thanks for sharing "
  }
}