{
  "id": 460191,
  "title": "4rd Place Solution for the Open Problems – Single-Cell Perturbations",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460191",
  "author_name": "paranoid",
  "post_date": "2023-12-08T06:33:26.238000",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>1. Context</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></li>\n</ul>\n<h1>2. Overview of the approach</h1>\n<p>Our approach is principally based on trial and error. Our optimal model is derived from several different models, of which the first model, RAPIDS SVR, was used to provide pseudo-labels but did not participate in the final ensembling:</p>\n<ol>\n<li>RAPIDS SVR - Used for training based on ChemEMB features to obtain pseudo-label results.</li>\n<li>Pyboost - Based on RAPIDS SVR pseudo-label features and other features. Public LB-0.581, Private LB-0.768.</li>\n<li>nn - Only utilized TruncatedSVD and Leaveoneout Encoder. Public LB-0.577, Private LB-0.731. (Yes, only nn can achieve 0.731 in Private LB)</li>\n<li>Open-source solution 0720 and open-source solution 0531 (late submission indicates that 0531 is not necessary) .<br>\nThe final solution's result is Public LB-0.566, Private LB-0.733~0.735. <br>\n<code>((0.3*Pyboost + 0.7*nn)*0.9 + open-source solution 0720*0.1)*0.95 + open-source solution 0531*0.05</code><br>\nThere is a score zone because our code has been saved in different versions in a messy manner, and there are some subtle parameter differences when reviewing our proposal, which makes it difficult to fully reproduce the final submission plan. This is our first time participating in the Kaggle competition, and in the future, we will be more meticulous in doing a good job of code version control.</li>\n</ol>\n<p>Different models employed different feature engineering strategies, which we will detail below.</p>\n<h1>3. Modeling</h1>\n<h2>3.1 Cross-validation</h2>\n<p>We believe that the method of cross-validation is a very important point for scoring improvement in this competition. A reasonable validation set allows us to trust our local CV instead of an overfitted LB. <br>\nInitially, we used random k-fold cross-validation. This method's CV maintained a consistent trend with the LB in the early stages. However, as we further incorporated nn models and added more training strategies, we noticed a discrepancy between the CV and LB. Consequently, we experimented with public two other forms of cross-validation which come from <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&amp;cellId=8\" target=\"_blank\">AmbrosM</a> and <a href=\"https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook\" target=\"_blank\">MT</a>. Thanks for their sharing. Ultimately, we chose MT's method of cross-validation, which showed excellent consistency between CV and LB in our models.</p>\n<h2>3.2 RAPIDS SVR</h2>\n<p>The first model we experimented with was RAPIDS SVR. It is frequently used in Kaggle competitions and has been part of multiple winning solutions. It is renowned for its fast training capabilities.</p>\n<h3>3.2.1 Feature Engineering</h3>\n<p>The feature engineering for the RAPIDS SVR model primarily included generating embedding features from ChemBERTa-10M-MTR for SMILES and statistical features obtained by aggregating target data. During the competition, we noticed in the discussion forums that many mentioned the embedding features generated by ChemBERTa-10M-MTR did not yield positive results. In our trials, we found that these features had a negative impact on models like LGBM and CatBoost, but they improved performance in RAPIDS SVR. Therefore, we decided to use RAPIDS SVR to generate pseudo-labels as a basis for subsequent model training.</p>\n<p>Here are the features we used in our RAPIDS SVR model:</p>\n<ol>\n<li>One-hot encoding features of cell_type and sm_name.</li>\n<li>Embedding features generated by ChemBERTa-10M-MTR. We used a mean pooling method comes from <a href=\"https://www.kaggle.com/code/cdeotte/rapids-svr-cv-0-450-lb-0-44x?scriptVersionId=105321484&amp;cellId=12\" target=\"_blank\">Chris</a> and set the maximum length for SMILES based on our analysis.</li>\n</ol>\n<pre><code>def mean_pooling(model_output, attention_mask):\n    token_embeddings = model_output.last_hidden_state.detach().cpu()\n    input_mask_expanded = (\n        attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()\n    )\n    return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(\n        input_mask_expanded.sum(1), =1e-9\n    )\n\nclass EmbedDataset(torch.utils.data.Dataset):\n    def __init__(self,df):\n        self.df = df.reset_index(=)\n    def __len__(self):\n        return len(self.df)\n    def __getitem__(self,idx):\n        text = self.df.loc[idx,]\n        tokens = tokenizer(\n                text,\n                None,\n                =,\n                =,\n                =,\n                =150,\n                =)\n        tokens = {k:v.squeeze(0)  k,v  tokens.items()}\n        return tokens\n\n\ndef get_embeddings(de_train, =, =150, =32, =):\n    global tokenizer\n    # Extract unique texts\n    unique_texts = de_train[].unique()\n\n    # Create a dataset  unique texts\n    ds_unique = EmbedDataset(pd.DataFrame(unique_texts, columns=[]))\n    embed_dataloader_unique = torch.utils.data.DataLoader(ds_unique, =BATCH_SIZE, =)\n\n    =\n    model = AutoModel.from_pretrained( MODEL_NM )\n    tokenizer = AutoTokenizer.from_pretrained( MODEL_NM )\n\n    model = model.(DEVICE)\n    model.eval()\n    unique_emb = []\n     batch  tqdm(embed_dataloader_unique,=len(embed_dataloader_unique)):\n        input_ids = batch[].(DEVICE)\n        attention_mask = batch[].(DEVICE)\n        with torch.no_grad():\n            model_output = model(=input_ids,attention_mask=attention_mask)\n        sentence_embeddings = mean_pooling(model_output, attention_mask.detach().cpu())\n        # Normalize the embeddings\n        sentence_embeddings = F.normalize(sentence_embeddings, =2, =1)\n        sentence_embeddings =  sentence_embeddings.squeeze(0).detach().cpu().numpy()\n        unique_emb.extend(sentence_embeddings)\n    unique_emb = np.array(unique_emb)\n     verbose:\n        (,unique_emb.shape)\n\n    text_to_embedding = {text: emb  text, emb  zip(unique_texts, unique_emb)}\n\n    train_emb = np.array([text_to_embedding[text]  text  de_train[]])\n    test_emb = np.array([text_to_embedding[text]  text  id_map[]])\n\n    return train_emb, test_emb\n\nMODEL_NM = \nall_train_text_feats, te_text_feats = get_embeddings(df_de_train, MODEL_NM)\n</code></pre>\n<p>3.Statistical features for cell_type and sm_name by aggregated from the target column. We retained features like ['mean', 'min', 'max', 'median', 'first', 'quantile_0.4'].</p>\n<p>Interestingly, when experimenting with different features, we did not train on all targets but selected only the first target, A1BG, for training and validation. This approach of feature selection allowed us to screen all features in just a few minutes, recording the CV scores for different features.</p>\n<h2>3.3 Pyboost</h2>\n<p>Our Pyboost solution is based on an open-source approach from <a href=\"https://www.kaggle.com/code/alexandervc/pyboost-secret-grandmaster-s-tool\" target=\"_blank\">Alexander Chervov</a>. Thanks for your sharing.</p>\n<h3>3.3.1 Feature Engineering</h3>\n<p>In our Pyboost model, we used four types of features:</p>\n<ol>\n<li>Pseudo-label features from RAPIDS SVR.</li>\n<li>Leaveoneout encoding features for cell_type and sm_name. We found in the Pyboost model that leaveoneout encoding was more effective than onehot encoding.</li>\n<li>Embedding features generated by ChemBERTa-10M-MTR for SMILES.</li>\n<li>Aggregated features for cell_type and sm_name against the target column, where we retained ['mean', 'max'].<br>\nWe reduced the dimensionality of 18,211 targets to 45 using TruncatedSVD. Similarly, we also reduced the dimensions of the above features to the same 45 dimensions. This dimensional reduction provided a certain improvement in our CV scores.</li>\n</ol>\n<h3>3.3.2 modeling</h3>\n<pre><code>model = GradientBoosting(\n                    \n                    ,=1000\n                    ,=0.01\n                    ,=10\n                    ,=1\n                    ,=0.2\n                    ,=1\n                    ,=0\n                    ,=100)\n</code></pre>\n<h2>3.4 nn</h2>\n<h3>3.4.1 Feature Engineering</h3>\n<p>In our nn model, we only used leaveoneout encoding features for cell_type and sm_name, as other features caused a decrease in CV scores. We also performed dimensionality reduction on the target data based on TruncatedSVD.</p>\n<h3>3.4.2 modeling</h3>\n<p>Our nn model consisted of a 3-layer 1D convolutional layer + 1 fully connected layer + a non-pretrained ResNet18 network. This structure allowed us to achieve a LB score of around 0.57 based solely on leaveoneout encoding features for cell_type and sm_name, which seems quite tricky. <br>\nOur initial nn model aimed to convert SMILES expressions into image data using the rdkit.Chem library and then input these images along with other features into the network for training, thus employing the ResNet network for image processing. However, we found that introducing SMILES image data did not improve training results. Despite this, we retained parts of the network structure and ended up with the aforementioned network.</p>\n<pre><code>class ResNetRegression(nn.Module):\n    def __init__(self, input_size, output_size, =, =32, =16, =0.2):\n        super(ResNetRegression, self).__init__()\n        self.reshape_size = reshape_size\n\n        self.conv1d_layers = nn.Sequential(\n            nn.Conv1d(1, num_channels, =3, =1, =1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n\n            nn.Conv1d(num_channels, num_channels, =3, =1, =1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n\n            nn.Conv1d(num_channels, num_channels, =3, =1, =1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n        )\n\n        self.fc_layers = nn.Sequential(\n            nn.Linear(num_channels * 4 * input_size, self.reshape_size * self.reshape_size),\n        )\n\n        self.resnet = models.resnet18(=pretrained)\n        self.resnet.conv1 = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), =)\n        self.resnet.fc = nn.Linear(self.resnet.fc.in_features, output_size)\n\n    def forward(self, x):\n        x = x.unsqueeze(1)  # Reshape x  Conv1d\n        x = self.conv1d_layers(x)\n        x = x.view(x.size(0), -1)  # Flatten  the linear layer\n        x = self.fc_layers(x)\n        x = x.view(x.size(0), 1, self.reshape_size, self.reshape_size)\n        x = self.resnet(x)\n        return x\n</code></pre>\n<h2>3.5 The 0720 Open-source Solution</h2>\n<p>In our final model, we assigned a weight of 0.05 to the 0720 open-source solution for ensembling. This is the result of a great notebook that uses the \"Autoencoder\" method. This integration resulted in an increase of 0.001 in our scores on both the Public LB and Private LB. We are grateful for the contribution shared by <a href=\"https://www.kaggle.com/vendekagonlabs\" target=\"_blank\">vendekagonlabs</a> and <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515\" target=\"_blank\">discussion</a>.</p>\n<h1>4 Parameter Tuning</h1>\n<p>We are not enthusiasts of parameter tuning, as, in our experience, tuning does not bring qualitative improvements to the model. In this competition, we only experimented with tuning towards the end, using Optuna to try out different settings for 'n_components' and 'n_iter' in TruncatedSVD, as well as 'sigma' in LeaveOneOutEncoder. Ultimately, we selected a few sets of parameters that yielded the best CV scores.</p>\n<h1>5 Things That Did Not Work</h1>\n<ol>\n<li>Normalization of target data.</li>\n<li>Converting SMILES into image features.</li>\n<li>Tree models such as LGBM and Catboost yielded average training results.</li>\n</ol>\n<p>Code<br>\n<a href=\"https://github.com/paralyzed2023/4st-place-solution-single-cell-pbs.git\" target=\"_blank\">https://github.com/paralyzed2023/4st-place-solution-single-cell-pbs.git</a></p>",
  "messages": [
    {
      "id": 2553318,
      "postDate": "2023-12-08T06:33:26.237Z",
      "content": "<h1>1. Context</h1>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data</a></li>\n</ul>\n<h1>2. Overview of the approach</h1>\n<p>Our approach is principally based on trial and error. Our optimal model is derived from several different models, of which the first model, RAPIDS SVR, was used to provide pseudo-labels but did not participate in the final ensembling:</p>\n<ol>\n<li>RAPIDS SVR - Used for training based on ChemEMB features to obtain pseudo-label results.</li>\n<li>Pyboost - Based on RAPIDS SVR pseudo-label features and other features. Public LB-0.581, Private LB-0.768.</li>\n<li>nn - Only utilized TruncatedSVD and Leaveoneout Encoder. Public LB-0.577, Private LB-0.731. (Yes, only nn can achieve 0.731 in Private LB)</li>\n<li>Open-source solution 0720 and open-source solution 0531 (late submission indicates that 0531 is not necessary) .<br>\nThe final solution's result is Public LB-0.566, Private LB-0.733~0.735. <br>\n<code>((0.3*Pyboost + 0.7*nn)*0.9 + open-source solution 0720*0.1)*0.95 + open-source solution 0531*0.05</code><br>\nThere is a score zone because our code has been saved in different versions in a messy manner, and there are some subtle parameter differences when reviewing our proposal, which makes it difficult to fully reproduce the final submission plan. This is our first time participating in the Kaggle competition, and in the future, we will be more meticulous in doing a good job of code version control.</li>\n</ol>\n<p>Different models employed different feature engineering strategies, which we will detail below.</p>\n<h1>3. Modeling</h1>\n<h2>3.1 Cross-validation</h2>\n<p>We believe that the method of cross-validation is a very important point for scoring improvement in this competition. A reasonable validation set allows us to trust our local CV instead of an overfitted LB. <br>\nInitially, we used random k-fold cross-validation. This method's CV maintained a consistent trend with the LB in the early stages. However, as we further incorporated nn models and added more training strategies, we noticed a discrepancy between the CV and LB. Consequently, we experimented with public two other forms of cross-validation which come from <a href=\"https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&amp;cellId=8\" target=\"_blank\">AmbrosM</a> and <a href=\"https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook\" target=\"_blank\">MT</a>. Thanks for their sharing. Ultimately, we chose MT's method of cross-validation, which showed excellent consistency between CV and LB in our models.</p>\n<h2>3.2 RAPIDS SVR</h2>\n<p>The first model we experimented with was RAPIDS SVR. It is frequently used in Kaggle competitions and has been part of multiple winning solutions. It is renowned for its fast training capabilities.</p>\n<h3>3.2.1 Feature Engineering</h3>\n<p>The feature engineering for the RAPIDS SVR model primarily included generating embedding features from ChemBERTa-10M-MTR for SMILES and statistical features obtained by aggregating target data. During the competition, we noticed in the discussion forums that many mentioned the embedding features generated by ChemBERTa-10M-MTR did not yield positive results. In our trials, we found that these features had a negative impact on models like LGBM and CatBoost, but they improved performance in RAPIDS SVR. Therefore, we decided to use RAPIDS SVR to generate pseudo-labels as a basis for subsequent model training.</p>\n<p>Here are the features we used in our RAPIDS SVR model:</p>\n<ol>\n<li>One-hot encoding features of cell_type and sm_name.</li>\n<li>Embedding features generated by ChemBERTa-10M-MTR. We used a mean pooling method comes from <a href=\"https://www.kaggle.com/code/cdeotte/rapids-svr-cv-0-450-lb-0-44x?scriptVersionId=105321484&amp;cellId=12\" target=\"_blank\">Chris</a> and set the maximum length for SMILES based on our analysis.</li>\n</ol>\n<pre><code>def mean_pooling(model_output, attention_mask):\n    token_embeddings = model_output.last_hidden_state.detach().cpu()\n    input_mask_expanded = (\n        attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()\n    )\n    return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(\n        input_mask_expanded.sum(1), =1e-9\n    )\n\nclass EmbedDataset(torch.utils.data.Dataset):\n    def __init__(self,df):\n        self.df = df.reset_index(=)\n    def __len__(self):\n        return len(self.df)\n    def __getitem__(self,idx):\n        text = self.df.loc[idx,]\n        tokens = tokenizer(\n                text,\n                None,\n                =,\n                =,\n                =,\n                =150,\n                =)\n        tokens = {k:v.squeeze(0)  k,v  tokens.items()}\n        return tokens\n\n\ndef get_embeddings(de_train, =, =150, =32, =):\n    global tokenizer\n    # Extract unique texts\n    unique_texts = de_train[].unique()\n\n    # Create a dataset  unique texts\n    ds_unique = EmbedDataset(pd.DataFrame(unique_texts, columns=[]))\n    embed_dataloader_unique = torch.utils.data.DataLoader(ds_unique, =BATCH_SIZE, =)\n\n    =\n    model = AutoModel.from_pretrained( MODEL_NM )\n    tokenizer = AutoTokenizer.from_pretrained( MODEL_NM )\n\n    model = model.(DEVICE)\n    model.eval()\n    unique_emb = []\n     batch  tqdm(embed_dataloader_unique,=len(embed_dataloader_unique)):\n        input_ids = batch[].(DEVICE)\n        attention_mask = batch[].(DEVICE)\n        with torch.no_grad():\n            model_output = model(=input_ids,attention_mask=attention_mask)\n        sentence_embeddings = mean_pooling(model_output, attention_mask.detach().cpu())\n        # Normalize the embeddings\n        sentence_embeddings = F.normalize(sentence_embeddings, =2, =1)\n        sentence_embeddings =  sentence_embeddings.squeeze(0).detach().cpu().numpy()\n        unique_emb.extend(sentence_embeddings)\n    unique_emb = np.array(unique_emb)\n     verbose:\n        (,unique_emb.shape)\n\n    text_to_embedding = {text: emb  text, emb  zip(unique_texts, unique_emb)}\n\n    train_emb = np.array([text_to_embedding[text]  text  de_train[]])\n    test_emb = np.array([text_to_embedding[text]  text  id_map[]])\n\n    return train_emb, test_emb\n\nMODEL_NM = \nall_train_text_feats, te_text_feats = get_embeddings(df_de_train, MODEL_NM)\n</code></pre>\n<p>3.Statistical features for cell_type and sm_name by aggregated from the target column. We retained features like ['mean', 'min', 'max', 'median', 'first', 'quantile_0.4'].</p>\n<p>Interestingly, when experimenting with different features, we did not train on all targets but selected only the first target, A1BG, for training and validation. This approach of feature selection allowed us to screen all features in just a few minutes, recording the CV scores for different features.</p>\n<h2>3.3 Pyboost</h2>\n<p>Our Pyboost solution is based on an open-source approach from <a href=\"https://www.kaggle.com/code/alexandervc/pyboost-secret-grandmaster-s-tool\" target=\"_blank\">Alexander Chervov</a>. Thanks for your sharing.</p>\n<h3>3.3.1 Feature Engineering</h3>\n<p>In our Pyboost model, we used four types of features:</p>\n<ol>\n<li>Pseudo-label features from RAPIDS SVR.</li>\n<li>Leaveoneout encoding features for cell_type and sm_name. We found in the Pyboost model that leaveoneout encoding was more effective than onehot encoding.</li>\n<li>Embedding features generated by ChemBERTa-10M-MTR for SMILES.</li>\n<li>Aggregated features for cell_type and sm_name against the target column, where we retained ['mean', 'max'].<br>\nWe reduced the dimensionality of 18,211 targets to 45 using TruncatedSVD. Similarly, we also reduced the dimensions of the above features to the same 45 dimensions. This dimensional reduction provided a certain improvement in our CV scores.</li>\n</ol>\n<h3>3.3.2 modeling</h3>\n<pre><code>model = GradientBoosting(\n                    \n                    ,=1000\n                    ,=0.01\n                    ,=10\n                    ,=1\n                    ,=0.2\n                    ,=1\n                    ,=0\n                    ,=100)\n</code></pre>\n<h2>3.4 nn</h2>\n<h3>3.4.1 Feature Engineering</h3>\n<p>In our nn model, we only used leaveoneout encoding features for cell_type and sm_name, as other features caused a decrease in CV scores. We also performed dimensionality reduction on the target data based on TruncatedSVD.</p>\n<h3>3.4.2 modeling</h3>\n<p>Our nn model consisted of a 3-layer 1D convolutional layer + 1 fully connected layer + a non-pretrained ResNet18 network. This structure allowed us to achieve a LB score of around 0.57 based solely on leaveoneout encoding features for cell_type and sm_name, which seems quite tricky. <br>\nOur initial nn model aimed to convert SMILES expressions into image data using the rdkit.Chem library and then input these images along with other features into the network for training, thus employing the ResNet network for image processing. However, we found that introducing SMILES image data did not improve training results. Despite this, we retained parts of the network structure and ended up with the aforementioned network.</p>\n<pre><code>class ResNetRegression(nn.Module):\n    def __init__(self, input_size, output_size, =, =32, =16, =0.2):\n        super(ResNetRegression, self).__init__()\n        self.reshape_size = reshape_size\n\n        self.conv1d_layers = nn.Sequential(\n            nn.Conv1d(1, num_channels, =3, =1, =1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n\n            nn.Conv1d(num_channels, num_channels, =3, =1, =1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n\n            nn.Conv1d(num_channels, num_channels, =3, =1, =1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n        )\n\n        self.fc_layers = nn.Sequential(\n            nn.Linear(num_channels * 4 * input_size, self.reshape_size * self.reshape_size),\n        )\n\n        self.resnet = models.resnet18(=pretrained)\n        self.resnet.conv1 = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), =)\n        self.resnet.fc = nn.Linear(self.resnet.fc.in_features, output_size)\n\n    def forward(self, x):\n        x = x.unsqueeze(1)  # Reshape x  Conv1d\n        x = self.conv1d_layers(x)\n        x = x.view(x.size(0), -1)  # Flatten  the linear layer\n        x = self.fc_layers(x)\n        x = x.view(x.size(0), 1, self.reshape_size, self.reshape_size)\n        x = self.resnet(x)\n        return x\n</code></pre>\n<h2>3.5 The 0720 Open-source Solution</h2>\n<p>In our final model, we assigned a weight of 0.05 to the 0720 open-source solution for ensembling. This is the result of a great notebook that uses the \"Autoencoder\" method. This integration resulted in an increase of 0.001 in our scores on both the Public LB and Private LB. We are grateful for the contribution shared by <a href=\"https://www.kaggle.com/vendekagonlabs\" target=\"_blank\">vendekagonlabs</a> and <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515\" target=\"_blank\">discussion</a>.</p>\n<h1>4 Parameter Tuning</h1>\n<p>We are not enthusiasts of parameter tuning, as, in our experience, tuning does not bring qualitative improvements to the model. In this competition, we only experimented with tuning towards the end, using Optuna to try out different settings for 'n_components' and 'n_iter' in TruncatedSVD, as well as 'sigma' in LeaveOneOutEncoder. Ultimately, we selected a few sets of parameters that yielded the best CV scores.</p>\n<h1>5 Things That Did Not Work</h1>\n<ol>\n<li>Normalization of target data.</li>\n<li>Converting SMILES into image features.</li>\n<li>Tree models such as LGBM and Catboost yielded average training results.</li>\n</ol>\n<p>Code<br>\n<a href=\"https://github.com/paralyzed2023/4st-place-solution-single-cell-pbs.git\" target=\"_blank\">https://github.com/paralyzed2023/4st-place-solution-single-cell-pbs.git</a></p>",
      "rawMarkdown": "# 1. Context\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n# 2. Overview of the approach\nOur approach is principally based on trial and error. Our optimal model is derived from several different models, of which the first model, RAPIDS SVR, was used to provide pseudo-labels but did not participate in the final ensembling:\n1. RAPIDS SVR - Used for training based on ChemEMB features to obtain pseudo-label results.\n2. Pyboost - Based on RAPIDS SVR pseudo-label features and other features. Public LB-0.581, Private LB-0.768.\n3. nn - Only utilized TruncatedSVD and Leaveoneout Encoder. Public LB-0.577, Private LB-0.731. (Yes, only nn can achieve 0.731 in Private LB)\n4. Open-source solution 0720 and open-source solution 0531 (late submission indicates that 0531 is not necessary) .\nThe final solution's result is Public LB-0.566, Private LB-0.733~0.735. \n`((0.3*Pyboost + 0.7*nn)*0.9 + open-source solution 0720*0.1)*0.95 + open-source solution 0531*0.05`\nThere is a score zone because our code has been saved in different versions in a messy manner, and there are some subtle parameter differences when reviewing our proposal, which makes it difficult to fully reproduce the final submission plan. This is our first time participating in the Kaggle competition, and in the future, we will be more meticulous in doing a good job of code version control.\n\nDifferent models employed different feature engineering strategies, which we will detail below.\n\n# 3. Modeling\n## 3.1 Cross-validation\nWe believe that the method of cross-validation is a very important point for scoring improvement in this competition. A reasonable validation set allows us to trust our local CV instead of an overfitted LB. \nInitially, we used random k-fold cross-validation. This method's CV maintained a consistent trend with the LB in the early stages. However, as we further incorporated nn models and added more training strategies, we noticed a discrepancy between the CV and LB. Consequently, we experimented with public two other forms of cross-validation which come from [AmbrosM](https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&cellId=8) and [MT](https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook). Thanks for their sharing. Ultimately, we chose MT's method of cross-validation, which showed excellent consistency between CV and LB in our models.\n\n## 3.2 RAPIDS SVR\nThe first model we experimented with was RAPIDS SVR. It is frequently used in Kaggle competitions and has been part of multiple winning solutions. It is renowned for its fast training capabilities.\n\n### 3.2.1 Feature Engineering\nThe feature engineering for the RAPIDS SVR model primarily included generating embedding features from ChemBERTa-10M-MTR for SMILES and statistical features obtained by aggregating target data. During the competition, we noticed in the discussion forums that many mentioned the embedding features generated by ChemBERTa-10M-MTR did not yield positive results. In our trials, we found that these features had a negative impact on models like LGBM and CatBoost, but they improved performance in RAPIDS SVR. Therefore, we decided to use RAPIDS SVR to generate pseudo-labels as a basis for subsequent model training.\n\nHere are the features we used in our RAPIDS SVR model:\n1. One-hot encoding features of cell_type and sm_name.\n2. Embedding features generated by ChemBERTa-10M-MTR. We used a mean pooling method comes from [Chris](https://www.kaggle.com/code/cdeotte/rapids-svr-cv-0-450-lb-0-44x?scriptVersionId=105321484&cellId=12) and set the maximum length for SMILES based on our analysis.\n```\ndef mean_pooling(model_output, attention_mask):\n    token_embeddings = model_output.last_hidden_state.detach().cpu()\n    input_mask_expanded = (\n        attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()\n    )\n    return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(\n        input_mask_expanded.sum(1), min=1e-9\n    )\n\nclass EmbedDataset(torch.utils.data.Dataset):\n    def __init__(self,df):\n        self.df = df.reset_index(drop=True)\n    def __len__(self):\n        return len(self.df)\n    def __getitem__(self,idx):\n        text = self.df.loc[idx,\"SMILES\"]\n        tokens = tokenizer(\n                text,\n                None,\n                add_special_tokens=True,\n                padding='max_length',\n                truncation=True,\n                max_length=150,\n                return_tensors=\"pt\")\n        tokens = {k:v.squeeze(0) for k,v in tokens.items()}\n        return tokens\n    \n\ndef get_embeddings(de_train, MODEL_NM='', MAX_LEN=150, BATCH_SIZE=32, verbose=True):\n    global tokenizer\n    # Extract unique texts\n    unique_texts = de_train[\"SMILES\"].unique()\n\n    # Create a dataset for unique texts\n    ds_unique = EmbedDataset(pd.DataFrame(unique_texts, columns=[\"SMILES\"]))\n    embed_dataloader_unique = torch.utils.data.DataLoader(ds_unique, batch_size=BATCH_SIZE, shuffle=False)\n\n    DEVICE=\"cuda\"\n    model = AutoModel.from_pretrained( MODEL_NM )\n    tokenizer = AutoTokenizer.from_pretrained( MODEL_NM )\n    \n    model = model.to(DEVICE)\n    model.eval()\n    unique_emb = []\n    for batch in tqdm(embed_dataloader_unique,total=len(embed_dataloader_unique)):\n        input_ids = batch[\"input_ids\"].to(DEVICE)\n        attention_mask = batch[\"attention_mask\"].to(DEVICE)\n        with torch.no_grad():\n            model_output = model(input_ids=input_ids,attention_mask=attention_mask)\n        sentence_embeddings = mean_pooling(model_output, attention_mask.detach().cpu())\n        # Normalize the embeddings\n        sentence_embeddings = F.normalize(sentence_embeddings, p=2, dim=1)\n        sentence_embeddings =  sentence_embeddings.squeeze(0).detach().cpu().numpy()\n        unique_emb.extend(sentence_embeddings)\n    unique_emb = np.array(unique_emb)\n    if verbose:\n        print('unique embeddings shape',unique_emb.shape)\n        \n    text_to_embedding = {text: emb for text, emb in zip(unique_texts, unique_emb)}\n\n    train_emb = np.array([text_to_embedding[text] for text in de_train['SMILES']])\n    test_emb = np.array([text_to_embedding[text] for text in id_map[\"SMILES\"]])\n        \n    return train_emb, test_emb\n\nMODEL_NM = 'DeepChem/ChemBERTa-10M-MTR'\nall_train_text_feats, te_text_feats = get_embeddings(df_de_train, MODEL_NM)\n```\n3.Statistical features for cell_type and sm_name by aggregated from the target column. We retained features like ['mean', 'min', 'max', 'median', 'first', 'quantile_0.4'].\n\nInterestingly, when experimenting with different features, we did not train on all targets but selected only the first target, A1BG, for training and validation. This approach of feature selection allowed us to screen all features in just a few minutes, recording the CV scores for different features.\n\n## 3.3 Pyboost\nOur Pyboost solution is based on an open-source approach from [Alexander Chervov](https://www.kaggle.com/code/alexandervc/pyboost-secret-grandmaster-s-tool). Thanks for your sharing.\n\n### 3.3.1 Feature Engineering\nIn our Pyboost model, we used four types of features:\n1. Pseudo-label features from RAPIDS SVR.\n2. Leaveoneout encoding features for cell_type and sm_name. We found in the Pyboost model that leaveoneout encoding was more effective than onehot encoding.\n3. Embedding features generated by ChemBERTa-10M-MTR for SMILES.\n4. Aggregated features for cell_type and sm_name against the target column, where we retained ['mean', 'max'].\nWe reduced the dimensionality of 18,211 targets to 45 using TruncatedSVD. Similarly, we also reduced the dimensions of the above features to the same 45 dimensions. This dimensional reduction provided a certain improvement in our CV scores.\n### 3.3.2 modeling\n```\nmodel = GradientBoosting(\n                    'mse'\n                    ,ntrees=1000\n                    ,lr=0.01\n                    ,max_depth=10\n                    ,subsample=1\n                    ,colsample=0.2\n                    ,min_data_in_leaf=1\n                    ,min_gain_to_split=0\n                    ,verbose=100)\n```\n## 3.4 nn\n### 3.4.1 Feature Engineering\nIn our nn model, we only used leaveoneout encoding features for cell_type and sm_name, as other features caused a decrease in CV scores. We also performed dimensionality reduction on the target data based on TruncatedSVD.\n\n### 3.4.2 modeling\nOur nn model consisted of a 3-layer 1D convolutional layer + 1 fully connected layer + a non-pretrained ResNet18 network. This structure allowed us to achieve a LB score of around 0.57 based solely on leaveoneout encoding features for cell_type and sm_name, which seems quite tricky. \nOur initial nn model aimed to convert SMILES expressions into image data using the rdkit.Chem library and then input these images along with other features into the network for training, thus employing the ResNet network for image processing. However, we found that introducing SMILES image data did not improve training results. Despite this, we retained parts of the network structure and ended up with the aforementioned network.\n```\nclass ResNetRegression(nn.Module):\n    def __init__(self, input_size, output_size, pretrained=False, reshape_size=32, num_channels=16, dropout_rate=0.2):\n        super(ResNetRegression, self).__init__()\n        self.reshape_size = reshape_size\n\n        self.conv1d_layers = nn.Sequential(\n            nn.Conv1d(1, num_channels, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n\n            nn.Conv1d(num_channels, num_channels*2, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n            \n            nn.Conv1d(num_channels*2, num_channels*4, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n        )\n\n        self.fc_layers = nn.Sequential(\n            nn.Linear(num_channels * 4 * input_size, self.reshape_size * self.reshape_size),\n        )\n\n        self.resnet = models.resnet18(pretrained=pretrained)\n        self.resnet.conv1 = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)\n        self.resnet.fc = nn.Linear(self.resnet.fc.in_features, output_size)\n\n    def forward(self, x):\n        x = x.unsqueeze(1)  # Reshape x for Conv1d\n        x = self.conv1d_layers(x)\n        x = x.view(x.size(0), -1)  # Flatten for the linear layer\n        x = self.fc_layers(x)\n        x = x.view(x.size(0), 1, self.reshape_size, self.reshape_size)\n        x = self.resnet(x)\n        return x\n```\n\n## 3.5 The 0720 Open-source Solution\nIn our final model, we assigned a weight of 0.05 to the 0720 open-source solution for ensembling. This is the result of a great notebook that uses the \"Autoencoder\" method. This integration resulted in an increase of 0.001 in our scores on both the Public LB and Private LB. We are grateful for the contribution shared by [vendekagonlabs](https://www.kaggle.com/vendekagonlabs) and [discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515).\n\n# 4 Parameter Tuning\nWe are not enthusiasts of parameter tuning, as, in our experience, tuning does not bring qualitative improvements to the model. In this competition, we only experimented with tuning towards the end, using Optuna to try out different settings for 'n_components' and 'n_iter' in TruncatedSVD, as well as 'sigma' in LeaveOneOutEncoder. Ultimately, we selected a few sets of parameters that yielded the best CV scores.\n\n# 5 Things That Did Not Work\n1. Normalization of target data.\n2. Converting SMILES into image features.\n3. Tree models such as LGBM and Catboost yielded average training results.\n\nCode\nhttps://github.com/paralyzed2023/4st-place-solution-single-cell-pbs.git",
      "votes": 9
    },
    {
      "id": 2553681,
      "postDate": "2023-12-08T12:42:56.883Z",
      "content": "<p>I'm scrolling down top solutions and let me tel you: I would never expect this competition be won with ResNet18 <br>\n🤩</p>",
      "rawMarkdown": "I'm scrolling down top solutions and let me tel you: I would never expect this competition be won with ResNet18 \n🤩",
      "votes": 1
    },
    {
      "id": 2559129,
      "postDate": "2023-12-12T16:11:02.917Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2553339,
      "postDate": "2023-12-08T07:00:20.377Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2553681,
      "author_name": "Pizzaboi",
      "author_url": "",
      "post_date": "2023-12-08T12:42:56.883000",
      "content": "<p>I'm scrolling down top solutions and let me tel you: I would never expect this competition be won with ResNet18 <br>\n🤩</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2559129,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-12T16:11:02.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2553339,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-12-08T07:00:20.377000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2553318": "# 1. Context\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\n- https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\n\n# 2. Overview of the approach\nOur approach is principally based on trial and error. Our optimal model is derived from several different models, of which the first model, RAPIDS SVR, was used to provide pseudo-labels but did not participate in the final ensembling:\n1. RAPIDS SVR - Used for training based on ChemEMB features to obtain pseudo-label results.\n2. Pyboost - Based on RAPIDS SVR pseudo-label features and other features. Public LB-0.581, Private LB-0.768.\n3. nn - Only utilized TruncatedSVD and Leaveoneout Encoder. Public LB-0.577, Private LB-0.731. (Yes, only nn can achieve 0.731 in Private LB)\n4. Open-source solution 0720 and open-source solution 0531 (late submission indicates that 0531 is not necessary) .\nThe final solution's result is Public LB-0.566, Private LB-0.733~0.735. \n`((0.3*Pyboost + 0.7*nn)*0.9 + open-source solution 0720*0.1)*0.95 + open-source solution 0531*0.05`\nThere is a score zone because our code has been saved in different versions in a messy manner, and there are some subtle parameter differences when reviewing our proposal, which makes it difficult to fully reproduce the final submission plan. This is our first time participating in the Kaggle competition, and in the future, we will be more meticulous in doing a good job of code version control.\n\nDifferent models employed different feature engineering strategies, which we will detail below.\n\n# 3. Modeling\n## 3.1 Cross-validation\nWe believe that the method of cross-validation is a very important point for scoring improvement in this competition. A reasonable validation set allows us to trust our local CV instead of an overfitted LB. \nInitially, we used random k-fold cross-validation. This method's CV maintained a consistent trend with the LB in the early stages. However, as we further incorporated nn models and added more training strategies, we noticed a discrepancy between the CV and LB. Consequently, we experimented with public two other forms of cross-validation which come from [AmbrosM](https://www.kaggle.com/code/ambrosm/scp-quickstart?scriptVersionId=144293041&cellId=8) and [MT](https://www.kaggle.com/code/masato114/scp-quickstart-another-cv-strategy/notebook). Thanks for their sharing. Ultimately, we chose MT's method of cross-validation, which showed excellent consistency between CV and LB in our models.\n\n## 3.2 RAPIDS SVR\nThe first model we experimented with was RAPIDS SVR. It is frequently used in Kaggle competitions and has been part of multiple winning solutions. It is renowned for its fast training capabilities.\n\n### 3.2.1 Feature Engineering\nThe feature engineering for the RAPIDS SVR model primarily included generating embedding features from ChemBERTa-10M-MTR for SMILES and statistical features obtained by aggregating target data. During the competition, we noticed in the discussion forums that many mentioned the embedding features generated by ChemBERTa-10M-MTR did not yield positive results. In our trials, we found that these features had a negative impact on models like LGBM and CatBoost, but they improved performance in RAPIDS SVR. Therefore, we decided to use RAPIDS SVR to generate pseudo-labels as a basis for subsequent model training.\n\nHere are the features we used in our RAPIDS SVR model:\n1. One-hot encoding features of cell_type and sm_name.\n2. Embedding features generated by ChemBERTa-10M-MTR. We used a mean pooling method comes from [Chris](https://www.kaggle.com/code/cdeotte/rapids-svr-cv-0-450-lb-0-44x?scriptVersionId=105321484&cellId=12) and set the maximum length for SMILES based on our analysis.\n```\ndef mean_pooling(model_output, attention_mask):\n    token_embeddings = model_output.last_hidden_state.detach().cpu()\n    input_mask_expanded = (\n        attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()\n    )\n    return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(\n        input_mask_expanded.sum(1), min=1e-9\n    )\n\nclass EmbedDataset(torch.utils.data.Dataset):\n    def __init__(self,df):\n        self.df = df.reset_index(drop=True)\n    def __len__(self):\n        return len(self.df)\n    def __getitem__(self,idx):\n        text = self.df.loc[idx,\"SMILES\"]\n        tokens = tokenizer(\n                text,\n                None,\n                add_special_tokens=True,\n                padding='max_length',\n                truncation=True,\n                max_length=150,\n                return_tensors=\"pt\")\n        tokens = {k:v.squeeze(0) for k,v in tokens.items()}\n        return tokens\n    \n\ndef get_embeddings(de_train, MODEL_NM='', MAX_LEN=150, BATCH_SIZE=32, verbose=True):\n    global tokenizer\n    # Extract unique texts\n    unique_texts = de_train[\"SMILES\"].unique()\n\n    # Create a dataset for unique texts\n    ds_unique = EmbedDataset(pd.DataFrame(unique_texts, columns=[\"SMILES\"]))\n    embed_dataloader_unique = torch.utils.data.DataLoader(ds_unique, batch_size=BATCH_SIZE, shuffle=False)\n\n    DEVICE=\"cuda\"\n    model = AutoModel.from_pretrained( MODEL_NM )\n    tokenizer = AutoTokenizer.from_pretrained( MODEL_NM )\n    \n    model = model.to(DEVICE)\n    model.eval()\n    unique_emb = []\n    for batch in tqdm(embed_dataloader_unique,total=len(embed_dataloader_unique)):\n        input_ids = batch[\"input_ids\"].to(DEVICE)\n        attention_mask = batch[\"attention_mask\"].to(DEVICE)\n        with torch.no_grad():\n            model_output = model(input_ids=input_ids,attention_mask=attention_mask)\n        sentence_embeddings = mean_pooling(model_output, attention_mask.detach().cpu())\n        # Normalize the embeddings\n        sentence_embeddings = F.normalize(sentence_embeddings, p=2, dim=1)\n        sentence_embeddings =  sentence_embeddings.squeeze(0).detach().cpu().numpy()\n        unique_emb.extend(sentence_embeddings)\n    unique_emb = np.array(unique_emb)\n    if verbose:\n        print('unique embeddings shape',unique_emb.shape)\n        \n    text_to_embedding = {text: emb for text, emb in zip(unique_texts, unique_emb)}\n\n    train_emb = np.array([text_to_embedding[text] for text in de_train['SMILES']])\n    test_emb = np.array([text_to_embedding[text] for text in id_map[\"SMILES\"]])\n        \n    return train_emb, test_emb\n\nMODEL_NM = 'DeepChem/ChemBERTa-10M-MTR'\nall_train_text_feats, te_text_feats = get_embeddings(df_de_train, MODEL_NM)\n```\n3.Statistical features for cell_type and sm_name by aggregated from the target column. We retained features like ['mean', 'min', 'max', 'median', 'first', 'quantile_0.4'].\n\nInterestingly, when experimenting with different features, we did not train on all targets but selected only the first target, A1BG, for training and validation. This approach of feature selection allowed us to screen all features in just a few minutes, recording the CV scores for different features.\n\n## 3.3 Pyboost\nOur Pyboost solution is based on an open-source approach from [Alexander Chervov](https://www.kaggle.com/code/alexandervc/pyboost-secret-grandmaster-s-tool). Thanks for your sharing.\n\n### 3.3.1 Feature Engineering\nIn our Pyboost model, we used four types of features:\n1. Pseudo-label features from RAPIDS SVR.\n2. Leaveoneout encoding features for cell_type and sm_name. We found in the Pyboost model that leaveoneout encoding was more effective than onehot encoding.\n3. Embedding features generated by ChemBERTa-10M-MTR for SMILES.\n4. Aggregated features for cell_type and sm_name against the target column, where we retained ['mean', 'max'].\nWe reduced the dimensionality of 18,211 targets to 45 using TruncatedSVD. Similarly, we also reduced the dimensions of the above features to the same 45 dimensions. This dimensional reduction provided a certain improvement in our CV scores.\n### 3.3.2 modeling\n```\nmodel = GradientBoosting(\n                    'mse'\n                    ,ntrees=1000\n                    ,lr=0.01\n                    ,max_depth=10\n                    ,subsample=1\n                    ,colsample=0.2\n                    ,min_data_in_leaf=1\n                    ,min_gain_to_split=0\n                    ,verbose=100)\n```\n## 3.4 nn\n### 3.4.1 Feature Engineering\nIn our nn model, we only used leaveoneout encoding features for cell_type and sm_name, as other features caused a decrease in CV scores. We also performed dimensionality reduction on the target data based on TruncatedSVD.\n\n### 3.4.2 modeling\nOur nn model consisted of a 3-layer 1D convolutional layer + 1 fully connected layer + a non-pretrained ResNet18 network. This structure allowed us to achieve a LB score of around 0.57 based solely on leaveoneout encoding features for cell_type and sm_name, which seems quite tricky. \nOur initial nn model aimed to convert SMILES expressions into image data using the rdkit.Chem library and then input these images along with other features into the network for training, thus employing the ResNet network for image processing. However, we found that introducing SMILES image data did not improve training results. Despite this, we retained parts of the network structure and ended up with the aforementioned network.\n```\nclass ResNetRegression(nn.Module):\n    def __init__(self, input_size, output_size, pretrained=False, reshape_size=32, num_channels=16, dropout_rate=0.2):\n        super(ResNetRegression, self).__init__()\n        self.reshape_size = reshape_size\n\n        self.conv1d_layers = nn.Sequential(\n            nn.Conv1d(1, num_channels, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n\n            nn.Conv1d(num_channels, num_channels*2, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n            \n            nn.Conv1d(num_channels*2, num_channels*4, kernel_size=3, stride=1, padding=1),\n            nn.ReLU(),\n            nn.Dropout(dropout_rate),\n        )\n\n        self.fc_layers = nn.Sequential(\n            nn.Linear(num_channels * 4 * input_size, self.reshape_size * self.reshape_size),\n        )\n\n        self.resnet = models.resnet18(pretrained=pretrained)\n        self.resnet.conv1 = nn.Conv2d(1, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False)\n        self.resnet.fc = nn.Linear(self.resnet.fc.in_features, output_size)\n\n    def forward(self, x):\n        x = x.unsqueeze(1)  # Reshape x for Conv1d\n        x = self.conv1d_layers(x)\n        x = x.view(x.size(0), -1)  # Flatten for the linear layer\n        x = self.fc_layers(x)\n        x = x.view(x.size(0), 1, self.reshape_size, self.reshape_size)\n        x = self.resnet(x)\n        return x\n```\n\n## 3.5 The 0720 Open-source Solution\nIn our final model, we assigned a weight of 0.05 to the 0720 open-source solution for ensembling. This is the result of a great notebook that uses the \"Autoencoder\" method. This integration resulted in an increase of 0.001 in our scores on both the Public LB and Private LB. We are grateful for the contribution shared by [vendekagonlabs](https://www.kaggle.com/vendekagonlabs) and [discussion](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/452515).\n\n# 4 Parameter Tuning\nWe are not enthusiasts of parameter tuning, as, in our experience, tuning does not bring qualitative improvements to the model. In this competition, we only experimented with tuning towards the end, using Optuna to try out different settings for 'n_components' and 'n_iter' in TruncatedSVD, as well as 'sigma' in LeaveOneOutEncoder. Ultimately, we selected a few sets of parameters that yielded the best CV scores.\n\n# 5 Things That Did Not Work\n1. Normalization of target data.\n2. Converting SMILES into image features.\n3. Tree models such as LGBM and Catboost yielded average training results.\n\nCode\nhttps://github.com/paralyzed2023/4st-place-solution-single-cell-pbs.git",
    "2553681": "I'm scrolling down top solutions and let me tel you: I would never expect this competition be won with ResNet18 \n🤩",
    "2559129": "",
    "2553339": ""
  }
}