{
  "id": 458764,
  "title": "SCP 21st Solution",
  "url": "/competitions/open-problems-single-cell-perturbations/writeups/suterusu-2-scp-21st-solution",
  "author_name": "",
  "post_date": "2023-12-12T02:33:38.233Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p><strong>Feature Engineering</strong></p>\n<p>Broadly two different sets of features were used for different models used in the final ensemble:</p>\n<p>Feature Set 1</p>\n<ol>\n<li>Cell type - one-hot encoded</li>\n<li>SMILES - converted to 2048 bit vectors using RDKit Morgan fingerprints</li>\n<li>Drug properties: for each SMILE, log P and log of Molar Refractivity, standard scaled</li>\n<li>Control - whether or not drug is a control (1) or not (0)</li>\n</ol>\n<p>Feature Set 2</p>\n<ol>\n<li>Same features as in Feature Set 1</li>\n<li>Average SVD embedding values for drug effects on all cell types AND average SVD embedding values for cell response to various drugs</li>\n<li>Exclusive of 2), or usage of average SVD embedding values for log fold-change for drug effects on all cell types AND average SVD emedding values for log fold-change for cell response to various drugs (either use 2 or 3)</li>\n</ol>\n<p>Note that the number of SVD singular values to keep was chosen according to the Gavish-Donohoe (GD) SVD hard threshold method: <a href=\"https://arxiv.org/abs/1305.5870\" target=\"_blank\">https://arxiv.org/abs/1305.5870</a></p>\n<p>Target</p>\n<ol>\n<li>GD SVD criterion used for singular value cutoff for low dimensional modes to keep</li>\n<li>SVD applied to target matrix and the singular value cutoff determined according to the Gavish-Donohoe threshold</li>\n<li>All models were trained against the SVD embedding of the original target, and model predictions were transformed back using the transpose of the V matrix</li>\n</ol>\n<p><strong>Model Architectures and Data Upsampling</strong></p>\n<p><strong>Model 1:</strong> </p>\n<p>Simple direct regression on SVD embedding targets</p>\n<ol>\n<li>8 layers Dense feed-forward neural netwok (5128 neurons per layer), output layer 114 neurons </li>\n<li>SELU activation each layer except for output layer (no activation, linear regression output)</li>\n<li>Output is in the SVD embedding space (114 columns)</li>\n<li>Loss: MAE or Pseudo-huber</li>\n<li>Epochs: 800</li>\n<li>Batch size: 16</li>\n<li>Cosine training schedule with warm restart every 200 epochs (alpha = 0.01, t_mul = 1.0, m_mul = 0.9)</li>\n<li>Stochastic weight averaging (SWA): SWA start from epoch 2</li>\n<li>Predictions for 18211 genes:<ol>\n<li>Let output be the predicted SVD embedding</li>\n<li>Take predicted SVD embedding and multiply by transpose of V matrix from SVD to get back to original 18211 representation</li>\n<li>Number of singular values to keep chosen according to Gavish-Donohoe threshold (see above)</li></ol></li>\n</ol>\n<p><strong>Model 2:</strong></p>\n<p>Same architecture as Model 1 however sample weights introduced to loss function.</p>\n<p>Sample weight scheme:</p>\n<ol>\n<li>From training set filter out drug-cell pairs where B cells / myeloid cells were exposed to same compounds</li>\n<li>From 1), exposure to same set of compounds but observed difference in target ( - log10(p_val) * sign(LFC)) should be attributable to cellular difference</li>\n<li>For each cell type not B cells / myeloid cells calculate a notion of \"distance\" from the filtered and observed targets for exposure under same drugs using a distance metric of choice, e.g. Frobenius norm of difference of target matrices</li>\n<li>For each cell type not B cells / myeloid cells average out this \"distance\" metric calculated in 3) and then subtract from 1 i.e. distance to B cell or myeloid cell would be 0 s.t. one minus this amount would give each cell type of prediction interest a score of 1, whilst cell types further away gets a lower score</li>\n<li>Divide each cell type by the minimum score of the 6 cell types as calculated in step 4), and use this number as a weight for each row in training based on cell type used in the experiment</li>\n<li>Model is trained on this weighted loss inclusive of each row's weight</li>\n</ol>\n<p><strong>Model 3:</strong></p>\n<p>Skip connections architecture:</p>\n<ol>\n<li>8 or 9 Dense layers</li>\n<li>Skip connections:<ol>\n<li>Input dimension: 2056</li>\n<li>Concatenate layer: Input concatenated with layer 2 pre-activation output (3072 neurons) leading to 5128 output dimension <br>\n(3072+2056) before feeding into SELU activation layer</li>\n<li>Additive skip connections: SELU output of concatenate layer (5128) + pre-activation output of layer 4 (5128), SELU<br>\noutput of layer 4 + pre-activation output of layer 6 (5128), SELU output of layer 6 + pre-activation output of layer 8<br>\n(5128 / used where network has 9 hidden layers)</li></ol></li>\n<li>Other details similar to Model 1</li>\n</ol>\n<p><strong>Model 4:</strong></p>\n<p>Model 1 architecture but using training error to identify hard to predict drug-cell pairs for upsampling. Upsampling was done by identifying index of training samples (rows) which were at or below at certain training error threshold and then amplified by making a new copies (integer multiples) of these rows to be concatenated to original training set. </p>\n<p>The thinking here was that since the problem for predicting interactions for B / myeloid cells is potentially underspecified and to be extrapolated from observed interactions of other cells, the drug-cell pairs that have high row-wise accuracy or low MAE (or other regression metric) are not as important and performance on these rows can be sacrificed for better performance on the rows in training which have low row-wise accuracy or low MAE (or other regression metric). The amplified set was also manually checked for inclusion of the small number of B / myeloid cell observations in training.</p>\n<p>Broadly three types of this upsampling procedure were used with various models</p>\n<p><strong>Upsampling procedure 1:</strong> Regression based row-wise metric (MAE) for determining cut-off threshold</p>\n<ol>\n<li>Simpler smaller neural network trained for 200 epochs on original training set</li>\n<li>Row-wise MAE computed for each sample</li>\n<li>Take median of 614 row-wise MAE metrics</li>\n<li>Take a positive multiple of this median (e.g. 3x or 15x) to select the base set of training rows to be upsampled</li>\n<li>Make K times more (e.g. 7x) copies of the training subset in 4) and concatenate to original training set</li>\n<li>Re-train larger model (could be any model architecture) on this upsampled training set</li>\n</ol>\n<p><strong>Upsampling procedure 2:</strong> Sign classification using logistic loss for determining cut-off threshold</p>\n<p>The thinking behind this approach is that sign may be important to get right as an individual prediction where the magnitude (-log10(p_val)) is correct but where sign is not is very consequential for RWRMSE metric.</p>\n<ol>\n<li><p>Same procedure as in prior upsampling procedure, except neural network with regression output is trained against the sign of<br>\nthe log fold-change (i.e. target matrix is composed of +1/-1)</p>\n<p>Logistic Loss = (1/n) * Sum(i from 1 to n) L(y, t) where<br>\nL(y, t) = ln(1 + exp(-y * t))</p>\n<p>With t being in {-1, +1} i.e. the sign of the log fold change</p></li>\n<li><p>Row-wise accuracy (%) is computed on the training set</p></li>\n<li><p>Choose a cutoff below which the training rows are to be upsampled. I used arbitrary cutoffs such as 75% or accuracy cutoffs<br>\n3 standard deviations below the mean row-wise accuracy</p></li>\n<li><p>Repeat upsampling procedure as in the previous procedure amplifying this subset an integer number of times and retrain a<br>\nlarger model on this exapnded training set</p></li>\n</ol>\n<p><strong>Upsampling procedure 3:</strong> Sign classification but focussed on rows with bad sign classification for small p-values</p>\n<p>Small p-values (e.g. less than 0.1) leads to large magnitudes when -log10 transformed, so intuition is get sign more correct for these as a bad sign classification flips these magnitudes to other side of real number line. </p>\n<p>Similar procedure to upsampling procedure, however we calculate accuracy only on subset of genes for each drug-cell pair where p-values are below a chosen threshold. Once these row-wise accuracy figures are computed, the same process as in the prior sign upsampling procedure is used to upsample a subset for retraining.</p>\n<p><strong>Model 5:</strong></p>\n<p>Triple regression head model with upsampling procedure and contrastive loss. The idea behind this architecture is to have share layers (e.g. 5 layers) between 3 different regression outputs. A \"contrastive\" loss (see below) was used to incentivise each regression head to learn a different hypothesis to the other 2 heads. This model architecture was mostly trained with sign upsampling procedure 2 as described in Model 4.</p>\n<p>Architecture:</p>\n<ol>\n<li><p>Shared weight layers: 5 Dense layers</p></li>\n<li><p>Activation: SELU for shared layers and regression heads, linear activation for regression outputs</p></li>\n<li><p>3 regression heads: [3072, 2048, 1024] neurons before output layer for training y_train SVD embeddings</p></li>\n<li><p>Contrastive loss: Sum(head 1 to 3) of regression loss for each head + contrast_weight * average_pairwise_dissimilarity</p>\n<p>K = n_head choose 2</p>\n<p>Average_Pairwise_Dissimilarity = 1/K * Sum(i from 1 to K) (Average Row-wise Cosine Similarity + 1.)</p>\n<p>If two non-zero vectors are exactly opposite, row-wise cosine similarity evaluates to -1. If they are exactly the same, we <br>\nget +1 and if they are orthogonal we get 0. Adding 1 to the average row-wise cosine similarity ensures the minimization<br>\nobjective goes to 0 (instead of -1).</p>\n<p>Contrastive loss essentially balances between each regression driving down bias but also learning distinctive hypotheses<br>\nfrom the data. The amount of contrast between the heads is controlled by the contrast_weight</p></li>\n<li><p>Each regression heads' output is multiplied by the transpose of the V matrix from SVD to get back predictions for original<br>\n18211 genes. </p></li>\n<li><p>Some submissions used the best head's predictions as determined by training error. Other predictions ensembled the 3 heads'<br>\npredictions by equal or training loss derived weights (lower loss -&gt; higher weight)</p></li>\n</ol>\n<p><strong>Final Submission</strong></p>\n<p>The two final submissions were LB RWRMSE weighted ensembles of the 16 best and 80 best submissions.</p>\n<p>For each submission I took the LB RWRMSE error, cubed them and subtracted from 1. to derive a score. These scores were then normalized against each other for the final weighted addition of the submissions.\\</p>\n<p><strong>Code</strong></p>\n<p><a href=\"https://github.com/maxleverage/kaggle-scp\" target=\"_blank\">https://github.com/maxleverage/kaggle-scp</a></p>",
  "messages": [
    {
      "id": "2545294",
      "postDate": "12/01/2023 12:17:08",
      "content": "<p><strong>Feature Engineering</strong></p>\n<p>Broadly two different sets of features were used for different models used in the final ensemble:</p>\n<p>Feature Set 1</p>\n<ol>\n<li>Cell type - one-hot encoded</li>\n<li>SMILES - converted to 2048 bit vectors using RDKit Morgan fingerprints</li>\n<li>Drug properties: for each SMILE, log P and log of Molar Refractivity, standard scaled</li>\n<li>Control - whether or not drug is a control (1) or not (0)</li>\n</ol>\n<p>Feature Set 2</p>\n<ol>\n<li>Same features as in Feature Set 1</li>\n<li>Average SVD embedding values for drug effects on all cell types AND average SVD embedding values for cell response to various drugs</li>\n<li>Exclusive of 2), or usage of average SVD embedding values for log fold-change for drug effects on all cell types AND average SVD emedding values for log fold-change for cell response to various drugs (either use 2 or 3)</li>\n</ol>\n<p>Note that the number of SVD singular values to keep was chosen according to the Gavish-Donohoe (GD) SVD hard threshold method: <a href=\"https://arxiv.org/abs/1305.5870\" target=\"_blank\">https://arxiv.org/abs/1305.5870</a></p>\n<p>Target</p>\n<ol>\n<li>GD SVD criterion used for singular value cutoff for low dimensional modes to keep</li>\n<li>SVD applied to target matrix and the singular value cutoff determined according to the Gavish-Donohoe threshold</li>\n<li>All models were trained against the SVD embedding of the original target, and model predictions were transformed back using the transpose of the V matrix</li>\n</ol>\n<p><strong>Model Architectures and Data Upsampling</strong></p>\n<p><strong>Model 1:</strong> </p>\n<p>Simple direct regression on SVD embedding targets</p>\n<ol>\n<li>8 layers Dense feed-forward neural netwok (5128 neurons per layer), output layer 114 neurons </li>\n<li>SELU activation each layer except for output layer (no activation, linear regression output)</li>\n<li>Output is in the SVD embedding space (114 columns)</li>\n<li>Loss: MAE or Pseudo-huber</li>\n<li>Epochs: 800</li>\n<li>Batch size: 16</li>\n<li>Cosine training schedule with warm restart every 200 epochs (alpha = 0.01, t_mul = 1.0, m_mul = 0.9)</li>\n<li>Stochastic weight averaging (SWA): SWA start from epoch 2</li>\n<li>Predictions for 18211 genes:<ol>\n<li>Let output be the predicted SVD embedding</li>\n<li>Take predicted SVD embedding and multiply by transpose of V matrix from SVD to get back to original 18211 representation</li>\n<li>Number of singular values to keep chosen according to Gavish-Donohoe threshold (see above)</li></ol></li>\n</ol>\n<p><strong>Model 2:</strong></p>\n<p>Same architecture as Model 1 however sample weights introduced to loss function.</p>\n<p>Sample weight scheme:</p>\n<ol>\n<li>From training set filter out drug-cell pairs where B cells / myeloid cells were exposed to same compounds</li>\n<li>From 1), exposure to same set of compounds but observed difference in target ( - log10(p_val) * sign(LFC)) should be attributable to cellular difference</li>\n<li>For each cell type not B cells / myeloid cells calculate a notion of \"distance\" from the filtered and observed targets for exposure under same drugs using a distance metric of choice, e.g. Frobenius norm of difference of target matrices</li>\n<li>For each cell type not B cells / myeloid cells average out this \"distance\" metric calculated in 3) and then subtract from 1 i.e. distance to B cell or myeloid cell would be 0 s.t. one minus this amount would give each cell type of prediction interest a score of 1, whilst cell types further away gets a lower score</li>\n<li>Divide each cell type by the minimum score of the 6 cell types as calculated in step 4), and use this number as a weight for each row in training based on cell type used in the experiment</li>\n<li>Model is trained on this weighted loss inclusive of each row's weight</li>\n</ol>\n<p><strong>Model 3:</strong></p>\n<p>Skip connections architecture:</p>\n<ol>\n<li>8 or 9 Dense layers</li>\n<li>Skip connections:<ol>\n<li>Input dimension: 2056</li>\n<li>Concatenate layer: Input concatenated with layer 2 pre-activation output (3072 neurons) leading to 5128 output dimension <br>\n(3072+2056) before feeding into SELU activation layer</li>\n<li>Additive skip connections: SELU output of concatenate layer (5128) + pre-activation output of layer 4 (5128), SELU<br>\noutput of layer 4 + pre-activation output of layer 6 (5128), SELU output of layer 6 + pre-activation output of layer 8<br>\n(5128 / used where network has 9 hidden layers)</li></ol></li>\n<li>Other details similar to Model 1</li>\n</ol>\n<p><strong>Model 4:</strong></p>\n<p>Model 1 architecture but using training error to identify hard to predict drug-cell pairs for upsampling. Upsampling was done by identifying index of training samples (rows) which were at or below at certain training error threshold and then amplified by making a new copies (integer multiples) of these rows to be concatenated to original training set. </p>\n<p>The thinking here was that since the problem for predicting interactions for B / myeloid cells is potentially underspecified and to be extrapolated from observed interactions of other cells, the drug-cell pairs that have high row-wise accuracy or low MAE (or other regression metric) are not as important and performance on these rows can be sacrificed for better performance on the rows in training which have low row-wise accuracy or low MAE (or other regression metric). The amplified set was also manually checked for inclusion of the small number of B / myeloid cell observations in training.</p>\n<p>Broadly three types of this upsampling procedure were used with various models</p>\n<p><strong>Upsampling procedure 1:</strong> Regression based row-wise metric (MAE) for determining cut-off threshold</p>\n<ol>\n<li>Simpler smaller neural network trained for 200 epochs on original training set</li>\n<li>Row-wise MAE computed for each sample</li>\n<li>Take median of 614 row-wise MAE metrics</li>\n<li>Take a positive multiple of this median (e.g. 3x or 15x) to select the base set of training rows to be upsampled</li>\n<li>Make K times more (e.g. 7x) copies of the training subset in 4) and concatenate to original training set</li>\n<li>Re-train larger model (could be any model architecture) on this upsampled training set</li>\n</ol>\n<p><strong>Upsampling procedure 2:</strong> Sign classification using logistic loss for determining cut-off threshold</p>\n<p>The thinking behind this approach is that sign may be important to get right as an individual prediction where the magnitude (-log10(p_val)) is correct but where sign is not is very consequential for RWRMSE metric.</p>\n<ol>\n<li><p>Same procedure as in prior upsampling procedure, except neural network with regression output is trained against the sign of<br>\nthe log fold-change (i.e. target matrix is composed of +1/-1)</p>\n<p>Logistic Loss = (1/n) * Sum(i from 1 to n) L(y, t) where<br>\nL(y, t) = ln(1 + exp(-y * t))</p>\n<p>With t being in {-1, +1} i.e. the sign of the log fold change</p></li>\n<li><p>Row-wise accuracy (%) is computed on the training set</p></li>\n<li><p>Choose a cutoff below which the training rows are to be upsampled. I used arbitrary cutoffs such as 75% or accuracy cutoffs<br>\n3 standard deviations below the mean row-wise accuracy</p></li>\n<li><p>Repeat upsampling procedure as in the previous procedure amplifying this subset an integer number of times and retrain a<br>\nlarger model on this exapnded training set</p></li>\n</ol>\n<p><strong>Upsampling procedure 3:</strong> Sign classification but focussed on rows with bad sign classification for small p-values</p>\n<p>Small p-values (e.g. less than 0.1) leads to large magnitudes when -log10 transformed, so intuition is get sign more correct for these as a bad sign classification flips these magnitudes to other side of real number line. </p>\n<p>Similar procedure to upsampling procedure, however we calculate accuracy only on subset of genes for each drug-cell pair where p-values are below a chosen threshold. Once these row-wise accuracy figures are computed, the same process as in the prior sign upsampling procedure is used to upsample a subset for retraining.</p>\n<p><strong>Model 5:</strong></p>\n<p>Triple regression head model with upsampling procedure and contrastive loss. The idea behind this architecture is to have share layers (e.g. 5 layers) between 3 different regression outputs. A \"contrastive\" loss (see below) was used to incentivise each regression head to learn a different hypothesis to the other 2 heads. This model architecture was mostly trained with sign upsampling procedure 2 as described in Model 4.</p>\n<p>Architecture:</p>\n<ol>\n<li><p>Shared weight layers: 5 Dense layers</p></li>\n<li><p>Activation: SELU for shared layers and regression heads, linear activation for regression outputs</p></li>\n<li><p>3 regression heads: [3072, 2048, 1024] neurons before output layer for training y_train SVD embeddings</p></li>\n<li><p>Contrastive loss: Sum(head 1 to 3) of regression loss for each head + contrast_weight * average_pairwise_dissimilarity</p>\n<p>K = n_head choose 2</p>\n<p>Average_Pairwise_Dissimilarity = 1/K * Sum(i from 1 to K) (Average Row-wise Cosine Similarity + 1.)</p>\n<p>If two non-zero vectors are exactly opposite, row-wise cosine similarity evaluates to -1. If they are exactly the same, we <br>\nget +1 and if they are orthogonal we get 0. Adding 1 to the average row-wise cosine similarity ensures the minimization<br>\nobjective goes to 0 (instead of -1).</p>\n<p>Contrastive loss essentially balances between each regression driving down bias but also learning distinctive hypotheses<br>\nfrom the data. The amount of contrast between the heads is controlled by the contrast_weight</p></li>\n<li><p>Each regression heads' output is multiplied by the transpose of the V matrix from SVD to get back predictions for original<br>\n18211 genes. </p></li>\n<li><p>Some submissions used the best head's predictions as determined by training error. Other predictions ensembled the 3 heads'<br>\npredictions by equal or training loss derived weights (lower loss -&gt; higher weight)</p></li>\n</ol>\n<p><strong>Final Submission</strong></p>\n<p>The two final submissions were LB RWRMSE weighted ensembles of the 16 best and 80 best submissions.</p>\n<p>For each submission I took the LB RWRMSE error, cubed them and subtracted from 1. to derive a score. These scores were then normalized against each other for the final weighted addition of the submissions.\\</p>\n<p><strong>Code</strong></p>\n<p><a href=\"https://github.com/maxleverage/kaggle-scp\" target=\"_blank\">https://github.com/maxleverage/kaggle-scp</a></p>",
      "rawMarkdown": "**Feature Engineering**\n\nBroadly two different sets of features were used for different models used in the final ensemble:\n\nFeature Set 1\n1. Cell type - one-hot encoded\n2. SMILES - converted to 2048 bit vectors using RDKit Morgan fingerprints\n3. Drug properties: for each SMILE, log P and log of Molar Refractivity, standard scaled\n4. Control - whether or not drug is a control (1) or not (0)\n\nFeature Set 2\n1. Same features as in Feature Set 1\n2. Average SVD embedding values for drug effects on all cell types AND average SVD embedding values for cell response to various drugs\n3. Exclusive of 2), or usage of average SVD embedding values for log fold-change for drug effects on all cell types AND average SVD emedding values for log fold-change for cell response to various drugs (either use 2 or 3)\n\nNote that the number of SVD singular values to keep was chosen according to the Gavish-Donohoe (GD) SVD hard threshold method: https://arxiv.org/abs/1305.5870\n\nTarget\n1. GD SVD criterion used for singular value cutoff for low dimensional modes to keep\n2. SVD applied to target matrix and the singular value cutoff determined according to the Gavish-Donohoe threshold\n3. All models were trained against the SVD embedding of the original target, and model predictions were transformed back using the transpose of the V matrix\n\n**Model Architectures and Data Upsampling**\n\n**Model 1:** \n\nSimple direct regression on SVD embedding targets\n1. 8 layers Dense feed-forward neural netwok (5128 neurons per layer), output layer 114 neurons \n2. SELU activation each layer except for output layer (no activation, linear regression output)\n3. Output is in the SVD embedding space (114 columns)\n4. Loss: MAE or Pseudo-huber\n5. Epochs: 800\n6. Batch size: 16\n7. Cosine training schedule with warm restart every 200 epochs (alpha = 0.01, t_mul = 1.0, m_mul = 0.9)\n8. Stochastic weight averaging (SWA): SWA start from epoch 2\n9. Predictions for 18211 genes:\n    1. Let output be the predicted SVD embedding\n    2. Take predicted SVD embedding and multiply by transpose of V matrix from SVD to get back to original 18211 representation\n    3. Number of singular values to keep chosen according to Gavish-Donohoe threshold (see above)\n\n**Model 2:**\n\nSame architecture as Model 1 however sample weights introduced to loss function.\n\nSample weight scheme:\n\n1. From training set filter out drug-cell pairs where B cells / myeloid cells were exposed to same compounds\n2. From 1), exposure to same set of compounds but observed difference in target ( - log10(p_val) * sign(LFC)) should be attributable to cellular difference\n3. For each cell type not B cells / myeloid cells calculate a notion of \"distance\" from the filtered and observed targets for exposure under same drugs using a distance metric of choice, e.g. Frobenius norm of difference of target matrices\n4. For each cell type not B cells / myeloid cells average out this \"distance\" metric calculated in 3) and then subtract from 1 i.e. distance to B cell or myeloid cell would be 0 s.t. one minus this amount would give each cell type of prediction interest a score of 1, whilst cell types further away gets a lower score\n5. Divide each cell type by the minimum score of the 6 cell types as calculated in step 4), and use this number as a weight for each row in training based on cell type used in the experiment\n6. Model is trained on this weighted loss inclusive of each row's weight\n\n**Model 3:**\n\nSkip connections architecture:\n\n1. 8 or 9 Dense layers\n2. Skip connections:\n    1. Input dimension: 2056\n    2. Concatenate layer: Input concatenated with layer 2 pre-activation output (3072 neurons) leading to 5128 output dimension \n    (3072+2056) before feeding into SELU activation layer\n    3. Additive skip connections: SELU output of concatenate layer (5128) + pre-activation output of layer 4 (5128), SELU\n    output of layer 4 + pre-activation output of layer 6 (5128), SELU output of layer 6 + pre-activation output of layer 8\n    (5128 / used where network has 9 hidden layers)\n3. Other details similar to Model 1\n\n**Model 4:**\n\nModel 1 architecture but using training error to identify hard to predict drug-cell pairs for upsampling. Upsampling was done by identifying index of training samples (rows) which were at or below at certain training error threshold and then amplified by making a new copies (integer multiples) of these rows to be concatenated to original training set. \n\nThe thinking here was that since the problem for predicting interactions for B / myeloid cells is potentially underspecified and to be extrapolated from observed interactions of other cells, the drug-cell pairs that have high row-wise accuracy or low MAE (or other regression metric) are not as important and performance on these rows can be sacrificed for better performance on the rows in training which have low row-wise accuracy or low MAE (or other regression metric). The amplified set was also manually checked for inclusion of the small number of B / myeloid cell observations in training.\n\nBroadly three types of this upsampling procedure were used with various models\n\n**Upsampling procedure 1:** Regression based row-wise metric (MAE) for determining cut-off threshold\n\n1. Simpler smaller neural network trained for 200 epochs on original training set\n2. Row-wise MAE computed for each sample\n3. Take median of 614 row-wise MAE metrics\n4. Take a positive multiple of this median (e.g. 3x or 15x) to select the base set of training rows to be upsampled\n5. Make K times more (e.g. 7x) copies of the training subset in 4) and concatenate to original training set\n6. Re-train larger model (could be any model architecture) on this upsampled training set\n\n**Upsampling procedure 2:** Sign classification using logistic loss for determining cut-off threshold\n\nThe thinking behind this approach is that sign may be important to get right as an individual prediction where the magnitude (-log10(p_val)) is correct but where sign is not is very consequential for RWRMSE metric.\n\n1. Same procedure as in prior upsampling procedure, except neural network with regression output is trained against the sign of\n   the log fold-change (i.e. target matrix is composed of +1/-1)\n\n    Logistic Loss = (1/n) * Sum(i from 1 to n) L(y, t) where\n    L(y, t) = ln(1 + exp(-y * t))\n    \n    With t being in {-1, +1} i.e. the sign of the log fold change\n\n2. Row-wise accuracy (%) is computed on the training set\n3. Choose a cutoff below which the training rows are to be upsampled. I used arbitrary cutoffs such as 75% or accuracy cutoffs\n   3 standard deviations below the mean row-wise accuracy\n4. Repeat upsampling procedure as in the previous procedure amplifying this subset an integer number of times and retrain a\n   larger model on this exapnded training set\n   \n**Upsampling procedure 3:** Sign classification but focussed on rows with bad sign classification for small p-values\n\nSmall p-values (e.g. less than 0.1) leads to large magnitudes when -log10 transformed, so intuition is get sign more correct for these as a bad sign classification flips these magnitudes to other side of real number line. \n\nSimilar procedure to upsampling procedure, however we calculate accuracy only on subset of genes for each drug-cell pair where p-values are below a chosen threshold. Once these row-wise accuracy figures are computed, the same process as in the prior sign upsampling procedure is used to upsample a subset for retraining.\n\n**Model 5:**\n\nTriple regression head model with upsampling procedure and contrastive loss. The idea behind this architecture is to have share layers (e.g. 5 layers) between 3 different regression outputs. A \"contrastive\" loss (see below) was used to incentivise each regression head to learn a different hypothesis to the other 2 heads. This model architecture was mostly trained with sign upsampling procedure 2 as described in Model 4.\n\nArchitecture:\n\n1. Shared weight layers: 5 Dense layers\n2. Activation: SELU for shared layers and regression heads, linear activation for regression outputs\n3. 3 regression heads: [3072, 2048, 1024] neurons before output layer for training y_train SVD embeddings\n4. Contrastive loss: Sum(head 1 to 3) of regression loss for each head + contrast_weight * average_pairwise_dissimilarity\n   \n   K = n_head choose 2\n   \n   Average_Pairwise_Dissimilarity = 1/K * Sum(i from 1 to K) (Average Row-wise Cosine Similarity + 1.)\n   \n   If two non-zero vectors are exactly opposite, row-wise cosine similarity evaluates to -1. If they are exactly the same, we \n   get +1 and if they are orthogonal we get 0. Adding 1 to the average row-wise cosine similarity ensures the minimization\n   objective goes to 0 (instead of -1).\n   \n   Contrastive loss essentially balances between each regression driving down bias but also learning distinctive hypotheses\n   from the data. The amount of contrast between the heads is controlled by the contrast_weight\n   \n5. Each regression heads' output is multiplied by the transpose of the V matrix from SVD to get back predictions for original\n   18211 genes. \n   \n6. Some submissions used the best head's predictions as determined by training error. Other predictions ensembled the 3 heads'\n   predictions by equal or training loss derived weights (lower loss -> higher weight)\n  \n**Final Submission**\n\nThe two final submissions were LB RWRMSE weighted ensembles of the 16 best and 80 best submissions.\n\nFor each submission I took the LB RWRMSE error, cubed them and subtracted from 1. to derive a score. These scores were then normalized against each other for the final weighted addition of the submissions.\\\n\n**Code**\n\nhttps://github.com/maxleverage/kaggle-scp",
      "votes": null
    },
    {
      "id": "2545498",
      "postDate": "12/01/2023 14:21:17",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/maxleverage\" target=\"_blank\">@maxleverage</a> 🎉 🎉 🎉, thanks for your sharing! Your method is excellent! But I have a question, as you mentioned in the text, are there any difference between Row-wise MAE and MAE?</p>",
      "rawMarkdown": "Congratulations! @maxleverage 🎉 🎉 🎉, thanks for your sharing! Your method is excellent! But I have a question, as you mentioned in the text, are there any difference between Row-wise MAE and MAE?",
      "votes": null
    },
    {
      "id": "2545968",
      "postDate": "12/02/2023 01:58:42",
      "content": "<p>No actually, the overall reduction to a scalar is the same (sum of elementwise MAE, divide by mxn)</p>",
      "rawMarkdown": "No actually, the overall reduction to a scalar is the same (sum of elementwise MAE, divide by mxn)",
      "votes": null
    },
    {
      "id": "2546592",
      "postDate": "12/02/2023 15:17:46",
      "content": "<p>Thank you! <a href=\"https://www.kaggle.com/maxleverage\" target=\"_blank\">@maxleverage</a> </p>",
      "rawMarkdown": "Thank you! @maxleverage",
      "votes": null
    },
    {
      "id": "2547454",
      "postDate": "12/03/2023 14:11:32",
      "content": "<p>But if ur using MAE to rows to upsample, then yes you would do the reduction across the columns (for each drug-cell pair, average the elementwise MAEs across all genes observed for that pair)</p>",
      "rawMarkdown": "But if ur using MAE to rows to upsample, then yes you would do the reduction across the columns (for each drug-cell pair, average the elementwise MAEs across all genes observed for that pair)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2545498,
      "author_name": "songqizhou",
      "author_url": "",
      "post_date": "12/01/2023 14:21:17",
      "content": "<p>Congratulations! <a href=\"https://www.kaggle.com/maxleverage\" target=\"_blank\">@maxleverage</a> 🎉 🎉 🎉, thanks for your sharing! Your method is excellent! But I have a question, as you mentioned in the text, are there any difference between Row-wise MAE and MAE?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2545968,
          "author_name": "maxleverage",
          "author_url": "",
          "post_date": "12/02/2023 01:58:42",
          "content": "<p>No actually, the overall reduction to a scalar is the same (sum of elementwise MAE, divide by mxn)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2546592,
              "author_name": "songqizhou",
              "author_url": "",
              "post_date": "12/02/2023 15:17:46",
              "content": "<p>Thank you! <a href=\"https://www.kaggle.com/maxleverage\" target=\"_blank\">@maxleverage</a> </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2547454,
                  "author_name": "maxleverage",
                  "author_url": "",
                  "post_date": "12/03/2023 14:11:32",
                  "content": "<p>But if ur using MAE to rows to upsample, then yes you would do the reduction across the columns (for each drug-cell pair, average the elementwise MAEs across all genes observed for that pair)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2545294": "**Feature Engineering**\n\nBroadly two different sets of features were used for different models used in the final ensemble:\n\nFeature Set 1\n1. Cell type - one-hot encoded\n2. SMILES - converted to 2048 bit vectors using RDKit Morgan fingerprints\n3. Drug properties: for each SMILE, log P and log of Molar Refractivity, standard scaled\n4. Control - whether or not drug is a control (1) or not (0)\n\nFeature Set 2\n1. Same features as in Feature Set 1\n2. Average SVD embedding values for drug effects on all cell types AND average SVD embedding values for cell response to various drugs\n3. Exclusive of 2), or usage of average SVD embedding values for log fold-change for drug effects on all cell types AND average SVD emedding values for log fold-change for cell response to various drugs (either use 2 or 3)\n\nNote that the number of SVD singular values to keep was chosen according to the Gavish-Donohoe (GD) SVD hard threshold method: https://arxiv.org/abs/1305.5870\n\nTarget\n1. GD SVD criterion used for singular value cutoff for low dimensional modes to keep\n2. SVD applied to target matrix and the singular value cutoff determined according to the Gavish-Donohoe threshold\n3. All models were trained against the SVD embedding of the original target, and model predictions were transformed back using the transpose of the V matrix\n\n**Model Architectures and Data Upsampling**\n\n**Model 1:** \n\nSimple direct regression on SVD embedding targets\n1. 8 layers Dense feed-forward neural netwok (5128 neurons per layer), output layer 114 neurons \n2. SELU activation each layer except for output layer (no activation, linear regression output)\n3. Output is in the SVD embedding space (114 columns)\n4. Loss: MAE or Pseudo-huber\n5. Epochs: 800\n6. Batch size: 16\n7. Cosine training schedule with warm restart every 200 epochs (alpha = 0.01, t_mul = 1.0, m_mul = 0.9)\n8. Stochastic weight averaging (SWA): SWA start from epoch 2\n9. Predictions for 18211 genes:\n    1. Let output be the predicted SVD embedding\n    2. Take predicted SVD embedding and multiply by transpose of V matrix from SVD to get back to original 18211 representation\n    3. Number of singular values to keep chosen according to Gavish-Donohoe threshold (see above)\n\n**Model 2:**\n\nSame architecture as Model 1 however sample weights introduced to loss function.\n\nSample weight scheme:\n\n1. From training set filter out drug-cell pairs where B cells / myeloid cells were exposed to same compounds\n2. From 1), exposure to same set of compounds but observed difference in target ( - log10(p_val) * sign(LFC)) should be attributable to cellular difference\n3. For each cell type not B cells / myeloid cells calculate a notion of \"distance\" from the filtered and observed targets for exposure under same drugs using a distance metric of choice, e.g. Frobenius norm of difference of target matrices\n4. For each cell type not B cells / myeloid cells average out this \"distance\" metric calculated in 3) and then subtract from 1 i.e. distance to B cell or myeloid cell would be 0 s.t. one minus this amount would give each cell type of prediction interest a score of 1, whilst cell types further away gets a lower score\n5. Divide each cell type by the minimum score of the 6 cell types as calculated in step 4), and use this number as a weight for each row in training based on cell type used in the experiment\n6. Model is trained on this weighted loss inclusive of each row's weight\n\n**Model 3:**\n\nSkip connections architecture:\n\n1. 8 or 9 Dense layers\n2. Skip connections:\n    1. Input dimension: 2056\n    2. Concatenate layer: Input concatenated with layer 2 pre-activation output (3072 neurons) leading to 5128 output dimension \n    (3072+2056) before feeding into SELU activation layer\n    3. Additive skip connections: SELU output of concatenate layer (5128) + pre-activation output of layer 4 (5128), SELU\n    output of layer 4 + pre-activation output of layer 6 (5128), SELU output of layer 6 + pre-activation output of layer 8\n    (5128 / used where network has 9 hidden layers)\n3. Other details similar to Model 1\n\n**Model 4:**\n\nModel 1 architecture but using training error to identify hard to predict drug-cell pairs for upsampling. Upsampling was done by identifying index of training samples (rows) which were at or below at certain training error threshold and then amplified by making a new copies (integer multiples) of these rows to be concatenated to original training set. \n\nThe thinking here was that since the problem for predicting interactions for B / myeloid cells is potentially underspecified and to be extrapolated from observed interactions of other cells, the drug-cell pairs that have high row-wise accuracy or low MAE (or other regression metric) are not as important and performance on these rows can be sacrificed for better performance on the rows in training which have low row-wise accuracy or low MAE (or other regression metric). The amplified set was also manually checked for inclusion of the small number of B / myeloid cell observations in training.\n\nBroadly three types of this upsampling procedure were used with various models\n\n**Upsampling procedure 1:** Regression based row-wise metric (MAE) for determining cut-off threshold\n\n1. Simpler smaller neural network trained for 200 epochs on original training set\n2. Row-wise MAE computed for each sample\n3. Take median of 614 row-wise MAE metrics\n4. Take a positive multiple of this median (e.g. 3x or 15x) to select the base set of training rows to be upsampled\n5. Make K times more (e.g. 7x) copies of the training subset in 4) and concatenate to original training set\n6. Re-train larger model (could be any model architecture) on this upsampled training set\n\n**Upsampling procedure 2:** Sign classification using logistic loss for determining cut-off threshold\n\nThe thinking behind this approach is that sign may be important to get right as an individual prediction where the magnitude (-log10(p_val)) is correct but where sign is not is very consequential for RWRMSE metric.\n\n1. Same procedure as in prior upsampling procedure, except neural network with regression output is trained against the sign of\n   the log fold-change (i.e. target matrix is composed of +1/-1)\n\n    Logistic Loss = (1/n) * Sum(i from 1 to n) L(y, t) where\n    L(y, t) = ln(1 + exp(-y * t))\n    \n    With t being in {-1, +1} i.e. the sign of the log fold change\n\n2. Row-wise accuracy (%) is computed on the training set\n3. Choose a cutoff below which the training rows are to be upsampled. I used arbitrary cutoffs such as 75% or accuracy cutoffs\n   3 standard deviations below the mean row-wise accuracy\n4. Repeat upsampling procedure as in the previous procedure amplifying this subset an integer number of times and retrain a\n   larger model on this exapnded training set\n   \n**Upsampling procedure 3:** Sign classification but focussed on rows with bad sign classification for small p-values\n\nSmall p-values (e.g. less than 0.1) leads to large magnitudes when -log10 transformed, so intuition is get sign more correct for these as a bad sign classification flips these magnitudes to other side of real number line. \n\nSimilar procedure to upsampling procedure, however we calculate accuracy only on subset of genes for each drug-cell pair where p-values are below a chosen threshold. Once these row-wise accuracy figures are computed, the same process as in the prior sign upsampling procedure is used to upsample a subset for retraining.\n\n**Model 5:**\n\nTriple regression head model with upsampling procedure and contrastive loss. The idea behind this architecture is to have share layers (e.g. 5 layers) between 3 different regression outputs. A \"contrastive\" loss (see below) was used to incentivise each regression head to learn a different hypothesis to the other 2 heads. This model architecture was mostly trained with sign upsampling procedure 2 as described in Model 4.\n\nArchitecture:\n\n1. Shared weight layers: 5 Dense layers\n2. Activation: SELU for shared layers and regression heads, linear activation for regression outputs\n3. 3 regression heads: [3072, 2048, 1024] neurons before output layer for training y_train SVD embeddings\n4. Contrastive loss: Sum(head 1 to 3) of regression loss for each head + contrast_weight * average_pairwise_dissimilarity\n   \n   K = n_head choose 2\n   \n   Average_Pairwise_Dissimilarity = 1/K * Sum(i from 1 to K) (Average Row-wise Cosine Similarity + 1.)\n   \n   If two non-zero vectors are exactly opposite, row-wise cosine similarity evaluates to -1. If they are exactly the same, we \n   get +1 and if they are orthogonal we get 0. Adding 1 to the average row-wise cosine similarity ensures the minimization\n   objective goes to 0 (instead of -1).\n   \n   Contrastive loss essentially balances between each regression driving down bias but also learning distinctive hypotheses\n   from the data. The amount of contrast between the heads is controlled by the contrast_weight\n   \n5. Each regression heads' output is multiplied by the transpose of the V matrix from SVD to get back predictions for original\n   18211 genes. \n   \n6. Some submissions used the best head's predictions as determined by training error. Other predictions ensembled the 3 heads'\n   predictions by equal or training loss derived weights (lower loss -> higher weight)\n  \n**Final Submission**\n\nThe two final submissions were LB RWRMSE weighted ensembles of the 16 best and 80 best submissions.\n\nFor each submission I took the LB RWRMSE error, cubed them and subtracted from 1. to derive a score. These scores were then normalized against each other for the final weighted addition of the submissions.\\\n\n**Code**\n\nhttps://github.com/maxleverage/kaggle-scp",
    "2545498": "Congratulations! @maxleverage 🎉 🎉 🎉, thanks for your sharing! Your method is excellent! But I have a question, as you mentioned in the text, are there any difference between Row-wise MAE and MAE?",
    "2545968": "No actually, the overall reduction to a scalar is the same (sum of elementwise MAE, divide by mxn)",
    "2546592": "Thank you! @maxleverage",
    "2547454": "But if ur using MAE to rows to upsample, then yes you would do the reduction across the columns (for each drug-cell pair, average the elementwise MAEs across all genes observed for that pair)"
  },
  "source": "meta"
}