{
  "id": 462031,
  "title": "17th Place Solution for the Open Problems – Single-Cell Perturbations (ADEFR)",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/462031",
  "author_name": "Ferenc Beres",
  "post_date": "2023-12-17T21:27:20.119000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>We are thrilled to share our solution with the community. Thanks to Kaggle and the Organizers for the opportunity to participate in such an interesting challenge.</p>\n<h2>1 Context</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Competition Overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Competition Data</a></li>\n</ul>\n<h2>2  Overview</h2>\n<ul>\n<li>Our solution comprises an ensemble of <strong>four gradient boosting based regression models</strong>, all trained on a shared feature set. Each model is endowed with distinctive hyperparameters and tailored training settings.</li>\n<li>Regression models are trained individually for each of the 129 compounds in the public/private test set, with the exception of instances where multi-regression is performed using <a href=\"https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb\" target=\"_blank\">SketchBoost</a>.</li>\n<li>Local training and evaluation follow a \"one-vs-rest\" style cross-validation approach across four cell types with known target variables: <code>['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells']</code>.</li>\n<li>Important features used:<ul>\n<li>Gene-wise PCA derived from the differential gene expression (DGE) table.</li>\n<li>Mean single-cell gene expression averaged over cell types and compounds, with particular emphasis on the control compound expression levels to enhance results.</li>\n<li>Gene-wise PCA derived from downsampled single-cell transcriptomics data.</li></ul></li>\n<li>We averaged feature values of nearest neighbor genes (neighbors were calculated based on the DGE data). This potentially reduced noise and improved downstream modeling.</li>\n<li>Moreover, we realized the impact of the number of cells used for DGE per <code>(cell type, compound)</code> on the mean target value per <code>(cell type, compound)</code>. Consequently, we incorporated the number of cells listed in the training data per <code>(cell type, compound)</code> to define importance weights for each training record using the formula: <code>w = 1-1/np.sqrt(1+num_cells['obs_id'][cell_type][compound])</code></li>\n<li>Our final submission is a blend of our solution (more details in the <code>4 Models</code> section) and a <a href=\"https://www.kaggle.com/code/olegpush/op2-eda-lb\" target=\"_blank\">public one</a> with 0.5-0.5 equal weights.</li>\n</ul>\n<h3>2.1 Things that did not  work</h3>\n<ul>\n<li>Inclusion of multiome data based features</li>\n<li>Inclusion of external knowledge based on<ul>\n<li>SMILES (we were experimenting with <a href=\"https://www.rdkit.org/docs/GettingStartedInPython.html\" target=\"_blank\">RDKit</a>)</li>\n<li>the <a href=\"https://string-db.org/\" target=\"_blank\">STRING database</a></li>\n<li><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7289078/\" target=\"_blank\">sci-Plex</a>, <a href=\"https://pubmed.ncbi.nlm.nih.gov/29195078/\" target=\"_blank\">LINCS data</a></li></ul></li>\n</ul>\n<h3>3 Feature engineering</h3>\n<h4>3.1 Input data</h4>\n<ul>\n<li>We extract features from the differential gene expression (DGE) data (<code>de_train.parquet</code>) and the single cell transcriptomics data (<code>adata_train.parquet</code>, <code>adata_obs_meta.csv</code>):</li>\n</ul>\n<pre><code> pandas  pd\n\n\ntr = pd.read_parquet(data_path + , engine=)\ngenes = (tr.columns[:])\nfeats = (tr[tr[]==][].unique())\n\n\ntrx = pd.read_pickle(data_path + )\ntro = pd.read_csv(data_path + )\ngenesx = (trx.index)\nobs = pd.DataFrame({:trx.columns}).reset_index().set_index()\nobs = obs.join(tro.set_index()).sort_values()\nX = trx.transpose()\nX = X.join(obs[[,]])\n...\n</code></pre>\n<ul>\n<li>We exclude single-cell records based on the file <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/454107\" target=\"_blank\">provided by Kaggle</a> </li>\n</ul>\n<pre><code>excluded_ids = pd.read_csv(data_path + )\nexcluded_ids[] = \nexcluded_ids = excluded_ids.pivot(index=, columns=, values=)\nexcluded_ids = excluded_ids.fillna()\nexcluded_ids = excluded_ids[trx.columns]\n\n = trx[[]]\n = .join(excluded_ids).fillna()\n =  (-)\n = np.array(.astype())\n\ntrx_ = np.array(trx)\ntrx_ =  * trx_\ntrx_ = pd.DataFrame(trx_)\ntrx_.columns = trx.columns\ntrx_.index = trx.index\ntrx = trx_\n...\n</code></pre>\n<ul>\n<li>Importantly, we downsample the single-cell data to get an <strong><em>equal number of cells per cell type and perturbation</em></strong> and we use this subsampled data for some of the downstream generated features</li>\n</ul>\n<pre><code>D = []\nsample_num = X.groupby([,]).count().().()\n(sample_num, end = )\n sm_name  X[].unique():\n  (sm_name,end = )\n   ct  X[].unique():\n    x = X[X[]==sm_name]\n    x = x[x[]==ct]\n    x = x[genesx]\n    x = x.head(sample_num)\n    D.append(x)  \nD = pd.concat(D,axis=)\n</code></pre>\n<h4>3.2 Single-cell transcriptomics data based PCA</h4>\n<ul>\n<li>We generate gene representations by performing PCA on this subsampled dataset</li>\n</ul>\n<pre><code> sklearn.decomposition  PCA\n\nn_components = \npca = PCA(n_components=n_components)\nd = D[genes].transpose()\n\nfeatures_pca_genes_sc = pca.fit_transform(d)\n...\n</code></pre>\n<h4>3.3 Mean gene expression based on the single-cell transcriptomics data</h4>\n<ul>\n<li>Mean gene expresison is calculated per cell type and compound</li>\n</ul>\n<pre><code>X = trx.transpose()\nX = X.join(obs[[,,]])\nX = X[X[]]\n X[]\nfeatures_mean_expression = {}\n cell_type  X[].unique():\n    (cell_type, end = )\n    features_mean_expression_ = X[X[]==cell_type]\n    features_mean_expression_ = features_mean_expression_.groupby()[genesx].mean()\n    ...\n</code></pre>\n<h4>3.4 DGE based PCA</h4>\n<pre><code>cell_types_train = [, , , ]\nD = tr[tr.apply( x: x[]  cell_types_train, axis = )][genes]\nD = D.transpose()\nD = (D - D.mean())/D.std()\nn_components = \npca = PCA(n_components=n_components)\nfeatures_pca_genes = pca.fit_transform(D)\n</code></pre>\n<h4>3.5 Raw DGE values of the 17 compounds known for all six cell types</h4>\n<pre><code>features_dge = {}\n cell_type  cell_types:\n    (cell_type, end = )\n    d = tr[tr[]==cell_type]\n    d = d.set_index()\n    d = d[genes].transpose()\n    features_ = pd.DataFrame()\n     feat  feats:\n         feat  d.columns:\n            features_[feat] = d[feat]\n        :\n            features_[feat] = \n    features_dge[cell_type] = features_\n...\n</code></pre>\n<h4>3.6 Nearest Neighbor based 'smoothing' over genes:</h4>\n<pre><code> sklearn.neighbors  NearestNeighbors\n\nfeats = (tr[tr[]==][].unique())\ntr_genes = tr[tr.apply( x: x[]  feats, axis = )][genes]\ntr_genes = tr_genes.transpose()\ntr_genes = (tr_genes - tr_genes.mean())/tr_genes.std()\ngene_nn = \nnbrs = NearestNeighbors(n_neighbors=gene_nn+).fit(tr_genes)\nnn_distances, nn_indices = nbrs.kneighbors(tr_genes)\n\n cell_type  cell_types:\n    (cell_type, end = )\n    d = features_dge[cell_type]\n    d_ = np.array(d)\n    d_nn = np.zeros(d.shape)\n     ii  ((d)):\n        n_neighbors = nn_indices.shape[]-\n        jj=\n         neighbor  nn_indices[ii,:]:\n            jj+=\n            d_nn[ii,:] += d_[neighbor,:]\n        d_nn[ii,:] = d_nn[ii,:]/n_neighbors\n    ...\n</code></pre>\n<h4>3.7 Smoothed DGE values:</h4>\n<p>We use Ridge regression to reduce the noise for DGE values</p>\n<pre><code> numpy  np\n\n ():\n    eps = eps * np.var(data, axis=)\n    nmat = data + np.random.normal(size=data.shape) * eps[, ...]\n    solvemat = np.linalg.inv(nmat.T.dot(nmat) + lambd * np.diag(np.ones(nmat.shape[]))).dot(nmat.T)\n    coefs = solvemat.dot(data)\n     coefs\n\nfeatures_dge_smoothed = {}\n cell_type  cell_types:\n    (cell_type, end = )\n    d = features_dge[cell_type]\n    d_ = np.array(d).transpose()\n    coefs = get_smoothing_coefs(d_)\n    d_smoothed = (d_.dot(coefs)).transpose()\n    d_smoothed = pd.DataFrame(d_smoothed)\n    d_smoothed.index = d.index\n    cols = []\n     col  feats:\n        cols.append(+col)\n    d_smoothed.columns = cols\n    features_dge_smoothed[cell_type] = d_smoothed\n</code></pre>\n<h3>4  Models</h3>\n<h4>4.1 Blending different boosting models</h4>\n<p>In our modeling framework, we use various boosting libraries with different settings.</p>\n<ul>\n<li><strong>LGBM:</strong> We predict DGE values for each of the 129 target compounds with different <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRegressor.html\" target=\"_blank\">LightGBM regressors</a>. This model achieved a public LB score of 0.571. </li>\n<li><strong>LGBM_LIN:</strong> We added the prediction of a simple Linear Regression as extra features to the previous setting. Public LB score is 0.573.</li>\n<li><strong>SketchBoost:</strong> We performed multi-regression for the 129 target compounds with a single <a href=\"[https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb]\" target=\"_blank\">SketchBoost</a> model. We trained this model with the following parameters: <code>{\"ntrees\":4096, \"lr\"=0.01, \"subsample\"=0.5, \"colsample\"=0.5, \"min_data_in_leaf\"=5, \"max_depth\"=5}</code></li>\n<li><strong>DART:</strong> We also use LightGBM with DART trained for 130 iterations with the following parameters: <code>{\"feature_fraction\": 0.7, \"bagging_fraction\": 0.6, \"data_sample_strategy\":'goss', 'boosting_type': 'dart'}</code>. Furthermore, we average out the predictions of five DART models trained with different random seeds.</li>\n</ul>\n<p>In our solution, we blend these models with equal (0.25) weights and then combine our prediction with a <a href=\"https://www.kaggle.com/code/olegpush/op2-eda-lb\" target=\"_blank\">public one</a>, as described above in the Overview section.</p>\n<h4>4.2. Feature importance</h4>\n<p>Below, we show some of our feature importance measurements for the SketchBoost component.<br>\nObserve that almost half of all importance is attributed to DGE-based gene PCA features (3.4). Ridge regression smoothing (3.7) proved to be slightly more effective than KNN-based smoothing (3.6). The remaining importance goes to raw compound DGE features (3.5) and mean expression values (3.3). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2Ff6b0a8cb4592370dd6a3a59d726fb68b%2Fcategory_imp.png?generation=1702845405654725&amp;alt=media\" alt=\"Top 20 most important features for SketchBoost\"></p>\n<p>Finally, we show the top 20 most important features for the SketchBoost component.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2F52ad6d64f8ab92a9cc0d1252b5c89399%2Ftop20.png?generation=1702845427700935&amp;alt=media\" alt=\"Top 20 most important features for SketchBoost\"></p>",
  "messages": [
    {
      "id": 2565271,
      "postDate": "2023-12-17T21:27:20.120Z",
      "content": "<p>We are thrilled to share our solution with the community. Thanks to Kaggle and the Organizers for the opportunity to participate in such an interesting challenge.</p>\n<h2>1 Context</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview\" target=\"_blank\">Competition Overview</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data\" target=\"_blank\">Competition Data</a></li>\n</ul>\n<h2>2  Overview</h2>\n<ul>\n<li>Our solution comprises an ensemble of <strong>four gradient boosting based regression models</strong>, all trained on a shared feature set. Each model is endowed with distinctive hyperparameters and tailored training settings.</li>\n<li>Regression models are trained individually for each of the 129 compounds in the public/private test set, with the exception of instances where multi-regression is performed using <a href=\"https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb\" target=\"_blank\">SketchBoost</a>.</li>\n<li>Local training and evaluation follow a \"one-vs-rest\" style cross-validation approach across four cell types with known target variables: <code>['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells']</code>.</li>\n<li>Important features used:<ul>\n<li>Gene-wise PCA derived from the differential gene expression (DGE) table.</li>\n<li>Mean single-cell gene expression averaged over cell types and compounds, with particular emphasis on the control compound expression levels to enhance results.</li>\n<li>Gene-wise PCA derived from downsampled single-cell transcriptomics data.</li></ul></li>\n<li>We averaged feature values of nearest neighbor genes (neighbors were calculated based on the DGE data). This potentially reduced noise and improved downstream modeling.</li>\n<li>Moreover, we realized the impact of the number of cells used for DGE per <code>(cell type, compound)</code> on the mean target value per <code>(cell type, compound)</code>. Consequently, we incorporated the number of cells listed in the training data per <code>(cell type, compound)</code> to define importance weights for each training record using the formula: <code>w = 1-1/np.sqrt(1+num_cells['obs_id'][cell_type][compound])</code></li>\n<li>Our final submission is a blend of our solution (more details in the <code>4 Models</code> section) and a <a href=\"https://www.kaggle.com/code/olegpush/op2-eda-lb\" target=\"_blank\">public one</a> with 0.5-0.5 equal weights.</li>\n</ul>\n<h3>2.1 Things that did not  work</h3>\n<ul>\n<li>Inclusion of multiome data based features</li>\n<li>Inclusion of external knowledge based on<ul>\n<li>SMILES (we were experimenting with <a href=\"https://www.rdkit.org/docs/GettingStartedInPython.html\" target=\"_blank\">RDKit</a>)</li>\n<li>the <a href=\"https://string-db.org/\" target=\"_blank\">STRING database</a></li>\n<li><a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7289078/\" target=\"_blank\">sci-Plex</a>, <a href=\"https://pubmed.ncbi.nlm.nih.gov/29195078/\" target=\"_blank\">LINCS data</a></li></ul></li>\n</ul>\n<h3>3 Feature engineering</h3>\n<h4>3.1 Input data</h4>\n<ul>\n<li>We extract features from the differential gene expression (DGE) data (<code>de_train.parquet</code>) and the single cell transcriptomics data (<code>adata_train.parquet</code>, <code>adata_obs_meta.csv</code>):</li>\n</ul>\n<pre><code> pandas  pd\n\n\ntr = pd.read_parquet(data_path + , engine=)\ngenes = (tr.columns[:])\nfeats = (tr[tr[]==][].unique())\n\n\ntrx = pd.read_pickle(data_path + )\ntro = pd.read_csv(data_path + )\ngenesx = (trx.index)\nobs = pd.DataFrame({:trx.columns}).reset_index().set_index()\nobs = obs.join(tro.set_index()).sort_values()\nX = trx.transpose()\nX = X.join(obs[[,]])\n...\n</code></pre>\n<ul>\n<li>We exclude single-cell records based on the file <a href=\"https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/454107\" target=\"_blank\">provided by Kaggle</a> </li>\n</ul>\n<pre><code>excluded_ids = pd.read_csv(data_path + )\nexcluded_ids[] = \nexcluded_ids = excluded_ids.pivot(index=, columns=, values=)\nexcluded_ids = excluded_ids.fillna()\nexcluded_ids = excluded_ids[trx.columns]\n\n = trx[[]]\n = .join(excluded_ids).fillna()\n =  (-)\n = np.array(.astype())\n\ntrx_ = np.array(trx)\ntrx_ =  * trx_\ntrx_ = pd.DataFrame(trx_)\ntrx_.columns = trx.columns\ntrx_.index = trx.index\ntrx = trx_\n...\n</code></pre>\n<ul>\n<li>Importantly, we downsample the single-cell data to get an <strong><em>equal number of cells per cell type and perturbation</em></strong> and we use this subsampled data for some of the downstream generated features</li>\n</ul>\n<pre><code>D = []\nsample_num = X.groupby([,]).count().().()\n(sample_num, end = )\n sm_name  X[].unique():\n  (sm_name,end = )\n   ct  X[].unique():\n    x = X[X[]==sm_name]\n    x = x[x[]==ct]\n    x = x[genesx]\n    x = x.head(sample_num)\n    D.append(x)  \nD = pd.concat(D,axis=)\n</code></pre>\n<h4>3.2 Single-cell transcriptomics data based PCA</h4>\n<ul>\n<li>We generate gene representations by performing PCA on this subsampled dataset</li>\n</ul>\n<pre><code> sklearn.decomposition  PCA\n\nn_components = \npca = PCA(n_components=n_components)\nd = D[genes].transpose()\n\nfeatures_pca_genes_sc = pca.fit_transform(d)\n...\n</code></pre>\n<h4>3.3 Mean gene expression based on the single-cell transcriptomics data</h4>\n<ul>\n<li>Mean gene expresison is calculated per cell type and compound</li>\n</ul>\n<pre><code>X = trx.transpose()\nX = X.join(obs[[,,]])\nX = X[X[]]\n X[]\nfeatures_mean_expression = {}\n cell_type  X[].unique():\n    (cell_type, end = )\n    features_mean_expression_ = X[X[]==cell_type]\n    features_mean_expression_ = features_mean_expression_.groupby()[genesx].mean()\n    ...\n</code></pre>\n<h4>3.4 DGE based PCA</h4>\n<pre><code>cell_types_train = [, , , ]\nD = tr[tr.apply( x: x[]  cell_types_train, axis = )][genes]\nD = D.transpose()\nD = (D - D.mean())/D.std()\nn_components = \npca = PCA(n_components=n_components)\nfeatures_pca_genes = pca.fit_transform(D)\n</code></pre>\n<h4>3.5 Raw DGE values of the 17 compounds known for all six cell types</h4>\n<pre><code>features_dge = {}\n cell_type  cell_types:\n    (cell_type, end = )\n    d = tr[tr[]==cell_type]\n    d = d.set_index()\n    d = d[genes].transpose()\n    features_ = pd.DataFrame()\n     feat  feats:\n         feat  d.columns:\n            features_[feat] = d[feat]\n        :\n            features_[feat] = \n    features_dge[cell_type] = features_\n...\n</code></pre>\n<h4>3.6 Nearest Neighbor based 'smoothing' over genes:</h4>\n<pre><code> sklearn.neighbors  NearestNeighbors\n\nfeats = (tr[tr[]==][].unique())\ntr_genes = tr[tr.apply( x: x[]  feats, axis = )][genes]\ntr_genes = tr_genes.transpose()\ntr_genes = (tr_genes - tr_genes.mean())/tr_genes.std()\ngene_nn = \nnbrs = NearestNeighbors(n_neighbors=gene_nn+).fit(tr_genes)\nnn_distances, nn_indices = nbrs.kneighbors(tr_genes)\n\n cell_type  cell_types:\n    (cell_type, end = )\n    d = features_dge[cell_type]\n    d_ = np.array(d)\n    d_nn = np.zeros(d.shape)\n     ii  ((d)):\n        n_neighbors = nn_indices.shape[]-\n        jj=\n         neighbor  nn_indices[ii,:]:\n            jj+=\n            d_nn[ii,:] += d_[neighbor,:]\n        d_nn[ii,:] = d_nn[ii,:]/n_neighbors\n    ...\n</code></pre>\n<h4>3.7 Smoothed DGE values:</h4>\n<p>We use Ridge regression to reduce the noise for DGE values</p>\n<pre><code> numpy  np\n\n ():\n    eps = eps * np.var(data, axis=)\n    nmat = data + np.random.normal(size=data.shape) * eps[, ...]\n    solvemat = np.linalg.inv(nmat.T.dot(nmat) + lambd * np.diag(np.ones(nmat.shape[]))).dot(nmat.T)\n    coefs = solvemat.dot(data)\n     coefs\n\nfeatures_dge_smoothed = {}\n cell_type  cell_types:\n    (cell_type, end = )\n    d = features_dge[cell_type]\n    d_ = np.array(d).transpose()\n    coefs = get_smoothing_coefs(d_)\n    d_smoothed = (d_.dot(coefs)).transpose()\n    d_smoothed = pd.DataFrame(d_smoothed)\n    d_smoothed.index = d.index\n    cols = []\n     col  feats:\n        cols.append(+col)\n    d_smoothed.columns = cols\n    features_dge_smoothed[cell_type] = d_smoothed\n</code></pre>\n<h3>4  Models</h3>\n<h4>4.1 Blending different boosting models</h4>\n<p>In our modeling framework, we use various boosting libraries with different settings.</p>\n<ul>\n<li><strong>LGBM:</strong> We predict DGE values for each of the 129 target compounds with different <a href=\"https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRegressor.html\" target=\"_blank\">LightGBM regressors</a>. This model achieved a public LB score of 0.571. </li>\n<li><strong>LGBM_LIN:</strong> We added the prediction of a simple Linear Regression as extra features to the previous setting. Public LB score is 0.573.</li>\n<li><strong>SketchBoost:</strong> We performed multi-regression for the 129 target compounds with a single <a href=\"[https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb]\" target=\"_blank\">SketchBoost</a> model. We trained this model with the following parameters: <code>{\"ntrees\":4096, \"lr\"=0.01, \"subsample\"=0.5, \"colsample\"=0.5, \"min_data_in_leaf\"=5, \"max_depth\"=5}</code></li>\n<li><strong>DART:</strong> We also use LightGBM with DART trained for 130 iterations with the following parameters: <code>{\"feature_fraction\": 0.7, \"bagging_fraction\": 0.6, \"data_sample_strategy\":'goss', 'boosting_type': 'dart'}</code>. Furthermore, we average out the predictions of five DART models trained with different random seeds.</li>\n</ul>\n<p>In our solution, we blend these models with equal (0.25) weights and then combine our prediction with a <a href=\"https://www.kaggle.com/code/olegpush/op2-eda-lb\" target=\"_blank\">public one</a>, as described above in the Overview section.</p>\n<h4>4.2. Feature importance</h4>\n<p>Below, we show some of our feature importance measurements for the SketchBoost component.<br>\nObserve that almost half of all importance is attributed to DGE-based gene PCA features (3.4). Ridge regression smoothing (3.7) proved to be slightly more effective than KNN-based smoothing (3.6). The remaining importance goes to raw compound DGE features (3.5) and mean expression values (3.3). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2Ff6b0a8cb4592370dd6a3a59d726fb68b%2Fcategory_imp.png?generation=1702845405654725&amp;alt=media\" alt=\"Top 20 most important features for SketchBoost\"></p>\n<p>Finally, we show the top 20 most important features for the SketchBoost component.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2F52ad6d64f8ab92a9cc0d1252b5c89399%2Ftop20.png?generation=1702845427700935&amp;alt=media\" alt=\"Top 20 most important features for SketchBoost\"></p>",
      "rawMarkdown": "We are thrilled to share our solution with the community. Thanks to Kaggle and the Organizers for the opportunity to participate in such an interesting challenge.\n\n## 1 Context\n\n- [Competition Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n- [Competition Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n## 2  Overview\n\n- Our solution comprises an ensemble of **four gradient boosting based regression models**, all trained on a shared feature set. Each model is endowed with distinctive hyperparameters and tailored training settings.\n- Regression models are trained individually for each of the 129 compounds in the public/private test set, with the exception of instances where multi-regression is performed using [SketchBoost](https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb).\n- Local training and evaluation follow a \"one-vs-rest\" style cross-validation approach across four cell types with known target variables: `['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells']`.\n- Important features used:\n  - Gene-wise PCA derived from the differential gene expression (DGE) table.\n  - Mean single-cell gene expression averaged over cell types and compounds, with particular emphasis on the control compound expression levels to enhance results.\n  - Gene-wise PCA derived from downsampled single-cell transcriptomics data.\n- We averaged feature values of nearest neighbor genes (neighbors were calculated based on the DGE data). This potentially reduced noise and improved downstream modeling.\n- Moreover, we realized the impact of the number of cells used for DGE per `(cell type, compound)` on the mean target value per `(cell type, compound)`. Consequently, we incorporated the number of cells listed in the training data per `(cell type, compound)` to define importance weights for each training record using the formula: `w = 1-1/np.sqrt(1+num_cells['obs_id'][cell_type][compound])`\n- Our final submission is a blend of our solution (more details in the `4 Models` section) and a [public one](https://www.kaggle.com/code/olegpush/op2-eda-lb) with 0.5-0.5 equal weights.\n\n### 2.1 Things that did not  work\n\n- Inclusion of multiome data based features\n- Inclusion of external knowledge based on\n  - SMILES (we were experimenting with [RDKit](https://www.rdkit.org/docs/GettingStartedInPython.html))\n  - the [STRING database](https://string-db.org/)\n  - [sci-Plex](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7289078/), [LINCS data](https://pubmed.ncbi.nlm.nih.gov/29195078/)\n\n### 3 Feature engineering\n\n#### 3.1 Input data\n\n- We extract features from the differential gene expression (DGE) data (`de_train.parquet`) and the single cell transcriptomics data (`adata_train.parquet`, `adata_obs_meta.csv`):\n\n```python\nimport pandas as pd\n\n# differential gene expression (DGE) data\ntr = pd.read_parquet(data_path + 'de_train.parquet', engine='pyarrow')\ngenes = list(tr.columns[5:])\nfeats = list(tr[tr['cell_type']=='B cells']['sm_name'].unique())\n\n# single cell data, pkl file includes normalized counts\ntrx = pd.read_pickle(data_path + 'adata_train.pkl')\ntro = pd.read_csv(data_path + 'adata_obs_meta.csv')\ngenesx = list(trx.index)\nobs = pd.DataFrame({'obs_id':trx.columns}).reset_index().set_index('obs_id')\nobs = obs.join(tro.set_index('obs_id')).sort_values('index')\nX = trx.transpose()\nX = X.join(obs[['cell_type','sm_name']])\n...\n```\n\n- We exclude single-cell records based on the file [provided by Kaggle](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/454107) \n\n```python\nexcluded_ids = pd.read_csv(data_path + 'adata_excluded_ids.csv')\nexcluded_ids['val'] = 1\nexcluded_ids = excluded_ids.pivot(index='gene', columns='obs_id', values='val')\nexcluded_ids = excluded_ids.fillna(0)\nexcluded_ids = excluded_ids[trx.columns]\n\nfilter = trx[[]]\nfilter = filter.join(excluded_ids).fillna(0)\nfilter =  (1-filter)\nfilter = np.array(filter.astype(bool))\n\ntrx_ = np.array(trx)\ntrx_ = filter * trx_\ntrx_ = pd.DataFrame(trx_)\ntrx_.columns = trx.columns\ntrx_.index = trx.index\ntrx = trx_\n...\n```\n\n- Importantly, we downsample the single-cell data to get an ***equal number of cells per cell type and perturbation*** and we use this subsampled data for some of the downstream generated features\n\n```python\nD = []\nsample_num = X.groupby(['cell_type','sm_name']).count().min().min()\nprint(sample_num, end = ' ')\nfor sm_name in X['sm_name'].unique():\n  print(sm_name,end = ' ')\n  for ct in X['cell_type'].unique():\n    x = X[X['sm_name']==sm_name]\n    x = x[x['cell_type']==ct]\n    x = x[genesx]\n    x = x.head(sample_num)\n    D.append(x)  \nD = pd.concat(D,axis=0)\n```\n\n#### 3.2 Single-cell transcriptomics data based PCA\n\n- We generate gene representations by performing PCA on this subsampled dataset\n\n```python\nfrom sklearn.decomposition import PCA\n\nn_components = 16\npca = PCA(n_components=n_components)\nd = D[genes].transpose()\n#d = (d - d.mean())/d.std()# in our experiments, normalization didn't improve performance\nfeatures_pca_genes_sc = pca.fit_transform(d)\n...\n```\n#### 3.3 Mean gene expression based on the single-cell transcriptomics data\n- Mean gene expresison is calculated per cell type and compound\n\n```python\nX = trx.transpose()\nX = X.join(obs[['control','cell_type','sm_name']])\nX = X[X['control']]\ndel X['control']\nfeatures_mean_expression = {}\nfor cell_type in X['cell_type'].unique():\n    print(cell_type, end = ' ')\n    features_mean_expression_ = X[X['cell_type']==cell_type]\n    features_mean_expression_ = features_mean_expression_.groupby('sm_name')[genesx].mean()\n    ...\n```\n\n#### 3.4 DGE based PCA\n\n```python\ncell_types_train = ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells']\nD = tr[tr.apply(lambda x: x['cell_type'] in cell_types_train, axis = 1)][genes]\nD = D.transpose()\nD = (D - D.mean())/D.std()\nn_components = 32\npca = PCA(n_components=n_components)\nfeatures_pca_genes = pca.fit_transform(D)\n```\n\n#### 3.5 Raw DGE values of the 17 compounds known for all six cell types\n\n```python\nfeatures_dge = {}\nfor cell_type in cell_types:\n    print(cell_type, end = ' ')\n    d = tr[tr['cell_type']==cell_type]\n    d = d.set_index('sm_name')\n    d = d[genes].transpose()\n    features_ = pd.DataFrame()\n    for feat in feats:\n        if feat in d.columns:\n            features_[feat] = d[feat]\n        else:\n            features_[feat] = 0\n    features_dge[cell_type] = features_\n...\n ```\n\n#### 3.6 Nearest Neighbor based 'smoothing' over genes:\n\n```python\nfrom sklearn.neighbors import NearestNeighbors\n\nfeats = list(tr[tr['cell_type']=='B cells']['sm_name'].unique())\ntr_genes = tr[tr.apply(lambda x: x['sm_name'] in feats, axis = 1)][genes]\ntr_genes = tr_genes.transpose()\ntr_genes = (tr_genes - tr_genes.mean())/tr_genes.std()\ngene_nn = 8\nnbrs = NearestNeighbors(n_neighbors=gene_nn+1).fit(tr_genes)\nnn_distances, nn_indices = nbrs.kneighbors(tr_genes)\n\nfor cell_type in cell_types:\n    print(cell_type, end = ' ')\n    d = features_dge[cell_type]\n    d_ = np.array(d)\n    d_nn = np.zeros(d.shape)\n    for ii in range(len(d)):\n        n_neighbors = nn_indices.shape[1]-1\n        jj=1\n        for neighbor in nn_indices[ii,1:]:\n            jj+=1\n            d_nn[ii,:] += d_[neighbor,:]\n        d_nn[ii,:] = d_nn[ii,:]/n_neighbors\n    ...\n```\n\n#### 3.7 Smoothed DGE values:\n\nWe use Ridge regression to reduce the noise for DGE values\n\n```python\nimport numpy as np\n\ndef get_smoothing_coefs(data, eps=0.01, lambd=10000):\n    eps = eps * np.var(data, axis=0)\n    nmat = data + np.random.normal(size=data.shape) * eps[None, ...]\n    solvemat = np.linalg.inv(nmat.T.dot(nmat) + lambd * np.diag(np.ones(nmat.shape[1]))).dot(nmat.T)\n    coefs = solvemat.dot(data)\n    return coefs\n\nfeatures_dge_smoothed = {}\nfor cell_type in cell_types:\n    print(cell_type, end = ' ')\n    d = features_dge[cell_type]\n    d_ = np.array(d).transpose()\n    coefs = get_smoothing_coefs(d_)\n    d_smoothed = (d_.dot(coefs)).transpose()\n    d_smoothed = pd.DataFrame(d_smoothed)\n    d_smoothed.index = d.index\n    cols = []\n    for col in feats:\n        cols.append('smoothed_'+col)\n    d_smoothed.columns = cols\n    features_dge_smoothed[cell_type] = d_smoothed\n```\n\n### 4  Models\n\n#### 4.1 Blending different boosting models\n\nIn our modeling framework, we use various boosting libraries with different settings.\n\n- **LGBM:** We predict DGE values for each of the 129 target compounds with different [LightGBM regressors](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRegressor.html). This model achieved a public LB score of 0.571. \n- **LGBM_LIN:** We added the prediction of a simple Linear Regression as extra features to the previous setting. Public LB score is 0.573.\n- **SketchBoost:** We performed multi-regression for the 129 target compounds with a single [SketchBoost]([https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb]) model. We trained this model with the following parameters: `{\"ntrees\":4096, \"lr\"=0.01, \"subsample\"=0.5, \"colsample\"=0.5, \"min_data_in_leaf\"=5, \"max_depth\"=5}`\n- **DART:** We also use LightGBM with DART trained for 130 iterations with the following parameters: `{\"feature_fraction\": 0.7, \"bagging_fraction\": 0.6, \"data_sample_strategy\":'goss', 'boosting_type': 'dart'}`. Furthermore, we average out the predictions of five DART models trained with different random seeds.\n\nIn our solution, we blend these models with equal (0.25) weights and then combine our prediction with a [public one](https://www.kaggle.com/code/olegpush/op2-eda-lb), as described above in the Overview section.\n\n#### 4.2. Feature importance\n\nBelow, we show some of our feature importance measurements for the SketchBoost component.\nObserve that almost half of all importance is attributed to DGE-based gene PCA features (3.4). Ridge regression smoothing (3.7) proved to be slightly more effective than KNN-based smoothing (3.6). The remaining importance goes to raw compound DGE features (3.5) and mean expression values (3.3). \n\n![Top 20 most important features for SketchBoost](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2Ff6b0a8cb4592370dd6a3a59d726fb68b%2Fcategory_imp.png?generation=1702845405654725&alt=media)\n\nFinally, we show the top 20 most important features for the SketchBoost component.\n\n![Top 20 most important features for SketchBoost](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2F52ad6d64f8ab92a9cc0d1252b5c89399%2Ftop20.png?generation=1702845427700935&alt=media)",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2565271": "We are thrilled to share our solution with the community. Thanks to Kaggle and the Organizers for the opportunity to participate in such an interesting challenge.\n\n## 1 Context\n\n- [Competition Overview](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/overview)\n- [Competition Data](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/data)\n\n## 2  Overview\n\n- Our solution comprises an ensemble of **four gradient boosting based regression models**, all trained on a shared feature set. Each model is endowed with distinctive hyperparameters and tailored training settings.\n- Regression models are trained individually for each of the 129 compounds in the public/private test set, with the exception of instances where multi-regression is performed using [SketchBoost](https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb).\n- Local training and evaluation follow a \"one-vs-rest\" style cross-validation approach across four cell types with known target variables: `['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells']`.\n- Important features used:\n  - Gene-wise PCA derived from the differential gene expression (DGE) table.\n  - Mean single-cell gene expression averaged over cell types and compounds, with particular emphasis on the control compound expression levels to enhance results.\n  - Gene-wise PCA derived from downsampled single-cell transcriptomics data.\n- We averaged feature values of nearest neighbor genes (neighbors were calculated based on the DGE data). This potentially reduced noise and improved downstream modeling.\n- Moreover, we realized the impact of the number of cells used for DGE per `(cell type, compound)` on the mean target value per `(cell type, compound)`. Consequently, we incorporated the number of cells listed in the training data per `(cell type, compound)` to define importance weights for each training record using the formula: `w = 1-1/np.sqrt(1+num_cells['obs_id'][cell_type][compound])`\n- Our final submission is a blend of our solution (more details in the `4 Models` section) and a [public one](https://www.kaggle.com/code/olegpush/op2-eda-lb) with 0.5-0.5 equal weights.\n\n### 2.1 Things that did not  work\n\n- Inclusion of multiome data based features\n- Inclusion of external knowledge based on\n  - SMILES (we were experimenting with [RDKit](https://www.rdkit.org/docs/GettingStartedInPython.html))\n  - the [STRING database](https://string-db.org/)\n  - [sci-Plex](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7289078/), [LINCS data](https://pubmed.ncbi.nlm.nih.gov/29195078/)\n\n### 3 Feature engineering\n\n#### 3.1 Input data\n\n- We extract features from the differential gene expression (DGE) data (`de_train.parquet`) and the single cell transcriptomics data (`adata_train.parquet`, `adata_obs_meta.csv`):\n\n```python\nimport pandas as pd\n\n# differential gene expression (DGE) data\ntr = pd.read_parquet(data_path + 'de_train.parquet', engine='pyarrow')\ngenes = list(tr.columns[5:])\nfeats = list(tr[tr['cell_type']=='B cells']['sm_name'].unique())\n\n# single cell data, pkl file includes normalized counts\ntrx = pd.read_pickle(data_path + 'adata_train.pkl')\ntro = pd.read_csv(data_path + 'adata_obs_meta.csv')\ngenesx = list(trx.index)\nobs = pd.DataFrame({'obs_id':trx.columns}).reset_index().set_index('obs_id')\nobs = obs.join(tro.set_index('obs_id')).sort_values('index')\nX = trx.transpose()\nX = X.join(obs[['cell_type','sm_name']])\n...\n```\n\n- We exclude single-cell records based on the file [provided by Kaggle](https://www.kaggle.com/competitions/open-problems-single-cell-perturbations/discussion/454107) \n\n```python\nexcluded_ids = pd.read_csv(data_path + 'adata_excluded_ids.csv')\nexcluded_ids['val'] = 1\nexcluded_ids = excluded_ids.pivot(index='gene', columns='obs_id', values='val')\nexcluded_ids = excluded_ids.fillna(0)\nexcluded_ids = excluded_ids[trx.columns]\n\nfilter = trx[[]]\nfilter = filter.join(excluded_ids).fillna(0)\nfilter =  (1-filter)\nfilter = np.array(filter.astype(bool))\n\ntrx_ = np.array(trx)\ntrx_ = filter * trx_\ntrx_ = pd.DataFrame(trx_)\ntrx_.columns = trx.columns\ntrx_.index = trx.index\ntrx = trx_\n...\n```\n\n- Importantly, we downsample the single-cell data to get an ***equal number of cells per cell type and perturbation*** and we use this subsampled data for some of the downstream generated features\n\n```python\nD = []\nsample_num = X.groupby(['cell_type','sm_name']).count().min().min()\nprint(sample_num, end = ' ')\nfor sm_name in X['sm_name'].unique():\n  print(sm_name,end = ' ')\n  for ct in X['cell_type'].unique():\n    x = X[X['sm_name']==sm_name]\n    x = x[x['cell_type']==ct]\n    x = x[genesx]\n    x = x.head(sample_num)\n    D.append(x)  \nD = pd.concat(D,axis=0)\n```\n\n#### 3.2 Single-cell transcriptomics data based PCA\n\n- We generate gene representations by performing PCA on this subsampled dataset\n\n```python\nfrom sklearn.decomposition import PCA\n\nn_components = 16\npca = PCA(n_components=n_components)\nd = D[genes].transpose()\n#d = (d - d.mean())/d.std()# in our experiments, normalization didn't improve performance\nfeatures_pca_genes_sc = pca.fit_transform(d)\n...\n```\n#### 3.3 Mean gene expression based on the single-cell transcriptomics data\n- Mean gene expresison is calculated per cell type and compound\n\n```python\nX = trx.transpose()\nX = X.join(obs[['control','cell_type','sm_name']])\nX = X[X['control']]\ndel X['control']\nfeatures_mean_expression = {}\nfor cell_type in X['cell_type'].unique():\n    print(cell_type, end = ' ')\n    features_mean_expression_ = X[X['cell_type']==cell_type]\n    features_mean_expression_ = features_mean_expression_.groupby('sm_name')[genesx].mean()\n    ...\n```\n\n#### 3.4 DGE based PCA\n\n```python\ncell_types_train = ['NK cells', 'T cells CD4+', 'T cells CD8+', 'T regulatory cells']\nD = tr[tr.apply(lambda x: x['cell_type'] in cell_types_train, axis = 1)][genes]\nD = D.transpose()\nD = (D - D.mean())/D.std()\nn_components = 32\npca = PCA(n_components=n_components)\nfeatures_pca_genes = pca.fit_transform(D)\n```\n\n#### 3.5 Raw DGE values of the 17 compounds known for all six cell types\n\n```python\nfeatures_dge = {}\nfor cell_type in cell_types:\n    print(cell_type, end = ' ')\n    d = tr[tr['cell_type']==cell_type]\n    d = d.set_index('sm_name')\n    d = d[genes].transpose()\n    features_ = pd.DataFrame()\n    for feat in feats:\n        if feat in d.columns:\n            features_[feat] = d[feat]\n        else:\n            features_[feat] = 0\n    features_dge[cell_type] = features_\n...\n ```\n\n#### 3.6 Nearest Neighbor based 'smoothing' over genes:\n\n```python\nfrom sklearn.neighbors import NearestNeighbors\n\nfeats = list(tr[tr['cell_type']=='B cells']['sm_name'].unique())\ntr_genes = tr[tr.apply(lambda x: x['sm_name'] in feats, axis = 1)][genes]\ntr_genes = tr_genes.transpose()\ntr_genes = (tr_genes - tr_genes.mean())/tr_genes.std()\ngene_nn = 8\nnbrs = NearestNeighbors(n_neighbors=gene_nn+1).fit(tr_genes)\nnn_distances, nn_indices = nbrs.kneighbors(tr_genes)\n\nfor cell_type in cell_types:\n    print(cell_type, end = ' ')\n    d = features_dge[cell_type]\n    d_ = np.array(d)\n    d_nn = np.zeros(d.shape)\n    for ii in range(len(d)):\n        n_neighbors = nn_indices.shape[1]-1\n        jj=1\n        for neighbor in nn_indices[ii,1:]:\n            jj+=1\n            d_nn[ii,:] += d_[neighbor,:]\n        d_nn[ii,:] = d_nn[ii,:]/n_neighbors\n    ...\n```\n\n#### 3.7 Smoothed DGE values:\n\nWe use Ridge regression to reduce the noise for DGE values\n\n```python\nimport numpy as np\n\ndef get_smoothing_coefs(data, eps=0.01, lambd=10000):\n    eps = eps * np.var(data, axis=0)\n    nmat = data + np.random.normal(size=data.shape) * eps[None, ...]\n    solvemat = np.linalg.inv(nmat.T.dot(nmat) + lambd * np.diag(np.ones(nmat.shape[1]))).dot(nmat.T)\n    coefs = solvemat.dot(data)\n    return coefs\n\nfeatures_dge_smoothed = {}\nfor cell_type in cell_types:\n    print(cell_type, end = ' ')\n    d = features_dge[cell_type]\n    d_ = np.array(d).transpose()\n    coefs = get_smoothing_coefs(d_)\n    d_smoothed = (d_.dot(coefs)).transpose()\n    d_smoothed = pd.DataFrame(d_smoothed)\n    d_smoothed.index = d.index\n    cols = []\n    for col in feats:\n        cols.append('smoothed_'+col)\n    d_smoothed.columns = cols\n    features_dge_smoothed[cell_type] = d_smoothed\n```\n\n### 4  Models\n\n#### 4.1 Blending different boosting models\n\nIn our modeling framework, we use various boosting libraries with different settings.\n\n- **LGBM:** We predict DGE values for each of the 129 target compounds with different [LightGBM regressors](https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.LGBMRegressor.html). This model achieved a public LB score of 0.571. \n- **LGBM_LIN:** We added the prediction of a simple Linear Regression as extra features to the previous setting. Public LB score is 0.573.\n- **SketchBoost:** We performed multi-regression for the 129 target compounds with a single [SketchBoost]([https://github.com/sb-ai-lab/Py-Boost/blob/master/tutorials/Tutorial_2_Advanced_multioutput.ipynb]) model. We trained this model with the following parameters: `{\"ntrees\":4096, \"lr\"=0.01, \"subsample\"=0.5, \"colsample\"=0.5, \"min_data_in_leaf\"=5, \"max_depth\"=5}`\n- **DART:** We also use LightGBM with DART trained for 130 iterations with the following parameters: `{\"feature_fraction\": 0.7, \"bagging_fraction\": 0.6, \"data_sample_strategy\":'goss', 'boosting_type': 'dart'}`. Furthermore, we average out the predictions of five DART models trained with different random seeds.\n\nIn our solution, we blend these models with equal (0.25) weights and then combine our prediction with a [public one](https://www.kaggle.com/code/olegpush/op2-eda-lb), as described above in the Overview section.\n\n#### 4.2. Feature importance\n\nBelow, we show some of our feature importance measurements for the SketchBoost component.\nObserve that almost half of all importance is attributed to DGE-based gene PCA features (3.4). Ridge regression smoothing (3.7) proved to be slightly more effective than KNN-based smoothing (3.6). The remaining importance goes to raw compound DGE features (3.5) and mean expression values (3.3). \n\n![Top 20 most important features for SketchBoost](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2Ff6b0a8cb4592370dd6a3a59d726fb68b%2Fcategory_imp.png?generation=1702845405654725&alt=media)\n\nFinally, we show the top 20 most important features for the SketchBoost component.\n\n![Top 20 most important features for SketchBoost](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F611446%2F52ad6d64f8ab92a9cc0d1252b5c89399%2Ftop20.png?generation=1702845427700935&alt=media)"
  }
}