{
  "id": 460059,
  "title": "A model used in the 24th solution - pure linear algebra!",
  "url": "/competitions/open-problems-single-cell-perturbations/discussion/460059",
  "author_name": "makio323",
  "post_date": "2023-12-07T17:18:27.487000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Intro</h1>\n<p>I share a linear algebra method, an unbiased and reproducible approach, resulting in decent Private/Public scores of 0.768/0.582.  My final submission is an ensemble of this <a href=\"https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582\" target=\"_blank\">Linear Algebra</a> approach along with AE NN mimicking the linear algebra approach by NN, whose joint weight is  0.70, and the combination of the public notebooks, <a href=\"https://www.kaggle.com/code/makio323/pyboost-secret-grandmaster-s-tool-0-592\" target=\"_blank\">Pyboost</a>, <a href=\"https://www.kaggle.com/code/makio323/fork-of-nlp-regression-12a31a-0-594\" target=\"_blank\">NN</a>, and <a href=\"https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain\" target=\"_blank\">Linear SVR</a>, with the total weight of 0.3.   It turns out that the pure Linear Algebra model gets the best private leaderboard score.</p>\n<p>In the first half of this competition, I struggled to overcome the wall of a 0.600 public score.  Some lucky runs of some NN models got over the wall, but not always.  There was also the second formidable wall of 0.585, which blended the results of public models.</p>\n<p>After the first half, I came up with this linear algebra approach and overcame the walls; the prediction is deterministic and reproducible and helped me to move on. </p>\n<p>The code of this linear model is available at <a href=\"https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582\" target=\"_blank\">my code notebook</a>. </p>\n<h1>Biological Hypothesis</h1>\n<p>It is a bit old, before the deep neural network, I had research experience using linear algebra to predict the missing values in a matrix -  <a href=\"https://www.cs.uic.edu/~mtamura/MakioTamuraMasterProject.pdf\" target=\"_blank\">Missing Value Expectation of Matrix Data by Fixed Rank Approximation Algorithm</a>, and it may inspire me.</p>\n<p>An assumption behind this method is that the differential expressions (DEs) of 18,211 genes at one cell line (e.g. NK cells) can be linearly transferable to those of another cell line (e.g. B Cells) on the same chemical perturbation.</p>\n<p>A chemical perturbation triggers a complex activity interaction among the 18,211 genes, and different chemical perturbations make different activity patterns, resulting in various DEs from the same baseline condition.  However, there would be an unseen “master rule” to govern these interactions on each cell line.  If the master rule of one cell line (e.g. NK cells) could be similar to another cell (e.g. B cells), DEs on the same chemical perturbation could be predictable from one to another.  Even not knowing the master rule of each cell line, their relationship among cell lines could be captured such that</p>\n<ul>\n<li><em>f</em>(DE<em>_i_c</em>) = DE<em>_j_c</em></li>\n</ul>\n<p>where DE<em>_i_x</em>and DE<em>_j_x</em> are the differential expressions of cell line <em>i</em> (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and <em>j</em> ('B cells', 'Myeloid cells') of chemical <em>c</em> perturbation, <em>f</em> is some special function.</p>\n<p>My approach is to assume that the linear system can be the proxy function, and to solve a system may provide the \"transfer\" such that</p>\n<ul>\n<li>DE<em>_i_core</em>  x T = DE<em>_j_core</em></li>\n</ul>\n<p>where DE<em>_i_core</em> and DE<em>_j_core</em> are <em>m</em> x <em>n</em> matrix, <em>m</em> is the number of chemicals shared by cel line <em>i</em> (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and <em>j</em> ('B cells', 'Myeloid cells') as shown below (including positive controls) and n is the number of genes (18,211).  T is an <em>n</em> x <em>n</em> matrix, considered a transformer matrix from one cell line to another.</p>\n<p>NK cells B cells 17<br>\n  NK cells Myeloid cells 17<br>\n  T cells CD4+ B cells 17<br>\n  T cells CD4+ Myeloid cells 17<br>\n  T cells CD8+ B cells 15<br>\n  T cells CD8+ Myeloid cells 15<br>\n  T regulatory cells B cells 17<br>\n  T regulatory cells Myeloid cells 17</p>\n<p>Once the transformer T is solved on the core chemicals, it may be applied to predict DEs of the target cell (e.g. B cells) from the known cell (e.g. NK cells)</p>\n<ul>\n<li>Prediction of DE<em>_i_target</em> = DE<em>_j_target</em>  x T</li>\n</ul>\n<p>where DE<em>_i_target</em> is the DEs of prediction cell line i (B cells and Myeloid cells) of the target chemicals, 128 for B cells and 127 for Myoloid cells, and DE<em>_j_target</em> is the DEs of known cell line j (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) of the target chemicals.</p>\n<p>This approach may provide a robust and unbiased prediction.</p>\n<h1>Observation</h1>\n<p>Using SVD to get the 1st and 2nd projected expression of the shared chemicals across 6 cell lines, and visualize the disperse among them (17 including positive controls, missing chemicals in T8 is replaced with those of T4).  Well, it is a bit difficult to recognize a clear pattern.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fb9f7d6b1595bae80513bbcd2fa786bb4%2Fs_plot.png?generation=1701966557440410&amp;alt=media\" alt=\"\"></p>\n<p>Here is a grid plot of the previous one by chemicals.  In most of the chemicals, there is not much difference among 6 cell cline.  However, NK cell seems to be similar to the B Cell and Myoloid cell on the chemical perturbations that create the larger difference among 6 cell lines such as Belinostat (one of the positive controls), MNL 2238, and Oprozomib.  In the end, accuracy in predicting such chemicals may be important. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fa5aac54f6b18ab37d2cdaa3e3c4c921e%2Fsg_plot.png?generation=1701966669627694&amp;alt=media\" alt=\"\"></p>\n<h1>Model</h1>\n<h3>Simpler model as in the background section.</h3>\n<ol>\n<li><p>Solve linear system, transformer, from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the DE using the chemicals tested in the two cell lines (15 plus 2 positive control).</p>\n<p>The transformer can be computed by multiplying a pseudo-inverse of DE<em>_i_core</em> with DE<em>_j_core</em> from the left.</p>\n<p>DE<em>_i_core</em>  x T = DE<em>_j_core</em><br>\nT = DE<em>_i_core</em>-t * DE<em>_j_core</em></p>\n<p>where DE<em>_i_core</em>-t is a pseudo-inverse of DE<em>_i_core</em> </p></li>\n<li><p>Apply the transformer to the DE of the base cell line/target chemicals and get the DE of the target cell line/chemicals.</p></li>\n</ol>\n<p>The simpler model consumes a lot of memory &gt; 20Gb, and it cannot be performed on the free version of the Saturn Cloud, which I had used for convenience, I also propose an alternative solution using SVD with a projection space.</p>\n<h3>SVD projection</h3>\n<p>It first reduces the gene dimension (18,211) to the full rank dimension of the entire data set (614) by SVD, create the transformer in the projected space, and applies the transformer in projected space from known cell to target cell on the target chemicals, reconstruct the original gene dimension of the prediction.</p>\n<ol>\n<li>Project the DE data by SVD with the whole dimension.</li>\n<li>Solve linear system from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the projected DE using the chemicals tested in the two cell lines (15 plus 2 positive control).</li>\n<li>Apply the transformer to the projected DE of the base cell line/target chemicals and get the projected DE of the target cell line/chemicals.</li>\n<li>Inverse the projected DE of the target cell line/chemicals to the predicted DE.</li>\n</ol>\n<h1>Robustness, Code, and Reproducibility</h1>\n<p>Please refer to <a href=\"https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582\" target=\"_blank\">my notebook</a>: the code can run on the notebook, and the result is deterministic, so reproducible. </p>\n<h1>Key Findings from the results</h1>\n<p>This method indeed can be used to diagnose the similarity of the \"master rule\" among the cell lines for scientific insight, and my result suggests that NK cells, along with T cells CD4+ can be stronger predictor cell lines for B Cell and Myeloid cells.   Interestingly, T cells CD4+ is a better predictor for B cell while NK cell is a better predictor for Myeloid cells.</p>\n<p>This finding may align with the biological findings - <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3072878/\" target=\"_blank\">NK cells can be derived from the myeloid lineage, Blood - The Journal of the American Society of Hematology 2011 3548 </a> and <a href=\"https://www.cell.com/trends/immunology/fulltext/S1471-4906(21)00117-4\" target=\"_blank\">T follicular helper cells cognately guide differentiation of antigen primed B cells in secondary lymphoid tissues - Trends in immunology 42.8 2021</a>.</p>\n<h3>Prediction by single base cell line</h3>\n<table>\n<thead>\n<tr>\n<th>base cell</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NK cells</td>\n<td>0.784</td>\n<td>0.596</td>\n</tr>\n<tr>\n<td>T cells CD4+</td>\n<td>0.775</td>\n<td>0.607</td>\n</tr>\n<tr>\n<td>T cells CD8+</td>\n<td>0.959</td>\n<td>0.706</td>\n</tr>\n<tr>\n<td>T regulatory cells</td>\n<td>0.834</td>\n<td>0.680</td>\n</tr>\n</tbody>\n</table>\n<h3>Prediction by two base cell lines</h3>\n<table>\n<thead>\n<tr>\n<th>base cell/target cell</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Predict B Cell by NK, Myeloid cells by T4</td>\n<td>0.786</td>\n<td>0.616</td>\n</tr>\n<tr>\n<td>Predict B Cell by T4, Myeloid cells by NK</td>\n<td>0.773</td>\n<td>0.587</td>\n</tr>\n</tbody>\n</table>\n<ol>\n<li><p>Differences in DE among certain cell lines can be captured linearly very well, and the linear relation is well transferable across different chemical responses, 0.768 in private and 0.582 in public</p></li>\n<li><p>DE of NK cell and T cells CD4+ cells are quite predictable for that of B Cell and Myeloid cells</p></li>\n<li><p>T cells CD4+ is a better predictor for B cell, and NK cells is a better predictor for Myeloid cell</p></li>\n</ol>",
  "messages": [
    {
      "id": 2552731,
      "postDate": "2023-12-07T17:18:27.487Z",
      "content": "<h1>Intro</h1>\n<p>I share a linear algebra method, an unbiased and reproducible approach, resulting in decent Private/Public scores of 0.768/0.582.  My final submission is an ensemble of this <a href=\"https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582\" target=\"_blank\">Linear Algebra</a> approach along with AE NN mimicking the linear algebra approach by NN, whose joint weight is  0.70, and the combination of the public notebooks, <a href=\"https://www.kaggle.com/code/makio323/pyboost-secret-grandmaster-s-tool-0-592\" target=\"_blank\">Pyboost</a>, <a href=\"https://www.kaggle.com/code/makio323/fork-of-nlp-regression-12a31a-0-594\" target=\"_blank\">NN</a>, and <a href=\"https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain\" target=\"_blank\">Linear SVR</a>, with the total weight of 0.3.   It turns out that the pure Linear Algebra model gets the best private leaderboard score.</p>\n<p>In the first half of this competition, I struggled to overcome the wall of a 0.600 public score.  Some lucky runs of some NN models got over the wall, but not always.  There was also the second formidable wall of 0.585, which blended the results of public models.</p>\n<p>After the first half, I came up with this linear algebra approach and overcame the walls; the prediction is deterministic and reproducible and helped me to move on. </p>\n<p>The code of this linear model is available at <a href=\"https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582\" target=\"_blank\">my code notebook</a>. </p>\n<h1>Biological Hypothesis</h1>\n<p>It is a bit old, before the deep neural network, I had research experience using linear algebra to predict the missing values in a matrix -  <a href=\"https://www.cs.uic.edu/~mtamura/MakioTamuraMasterProject.pdf\" target=\"_blank\">Missing Value Expectation of Matrix Data by Fixed Rank Approximation Algorithm</a>, and it may inspire me.</p>\n<p>An assumption behind this method is that the differential expressions (DEs) of 18,211 genes at one cell line (e.g. NK cells) can be linearly transferable to those of another cell line (e.g. B Cells) on the same chemical perturbation.</p>\n<p>A chemical perturbation triggers a complex activity interaction among the 18,211 genes, and different chemical perturbations make different activity patterns, resulting in various DEs from the same baseline condition.  However, there would be an unseen “master rule” to govern these interactions on each cell line.  If the master rule of one cell line (e.g. NK cells) could be similar to another cell (e.g. B cells), DEs on the same chemical perturbation could be predictable from one to another.  Even not knowing the master rule of each cell line, their relationship among cell lines could be captured such that</p>\n<ul>\n<li><em>f</em>(DE<em>_i_c</em>) = DE<em>_j_c</em></li>\n</ul>\n<p>where DE<em>_i_x</em>and DE<em>_j_x</em> are the differential expressions of cell line <em>i</em> (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and <em>j</em> ('B cells', 'Myeloid cells') of chemical <em>c</em> perturbation, <em>f</em> is some special function.</p>\n<p>My approach is to assume that the linear system can be the proxy function, and to solve a system may provide the \"transfer\" such that</p>\n<ul>\n<li>DE<em>_i_core</em>  x T = DE<em>_j_core</em></li>\n</ul>\n<p>where DE<em>_i_core</em> and DE<em>_j_core</em> are <em>m</em> x <em>n</em> matrix, <em>m</em> is the number of chemicals shared by cel line <em>i</em> (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and <em>j</em> ('B cells', 'Myeloid cells') as shown below (including positive controls) and n is the number of genes (18,211).  T is an <em>n</em> x <em>n</em> matrix, considered a transformer matrix from one cell line to another.</p>\n<p>NK cells B cells 17<br>\n  NK cells Myeloid cells 17<br>\n  T cells CD4+ B cells 17<br>\n  T cells CD4+ Myeloid cells 17<br>\n  T cells CD8+ B cells 15<br>\n  T cells CD8+ Myeloid cells 15<br>\n  T regulatory cells B cells 17<br>\n  T regulatory cells Myeloid cells 17</p>\n<p>Once the transformer T is solved on the core chemicals, it may be applied to predict DEs of the target cell (e.g. B cells) from the known cell (e.g. NK cells)</p>\n<ul>\n<li>Prediction of DE<em>_i_target</em> = DE<em>_j_target</em>  x T</li>\n</ul>\n<p>where DE<em>_i_target</em> is the DEs of prediction cell line i (B cells and Myeloid cells) of the target chemicals, 128 for B cells and 127 for Myoloid cells, and DE<em>_j_target</em> is the DEs of known cell line j (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) of the target chemicals.</p>\n<p>This approach may provide a robust and unbiased prediction.</p>\n<h1>Observation</h1>\n<p>Using SVD to get the 1st and 2nd projected expression of the shared chemicals across 6 cell lines, and visualize the disperse among them (17 including positive controls, missing chemicals in T8 is replaced with those of T4).  Well, it is a bit difficult to recognize a clear pattern.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fb9f7d6b1595bae80513bbcd2fa786bb4%2Fs_plot.png?generation=1701966557440410&amp;alt=media\" alt=\"\"></p>\n<p>Here is a grid plot of the previous one by chemicals.  In most of the chemicals, there is not much difference among 6 cell cline.  However, NK cell seems to be similar to the B Cell and Myoloid cell on the chemical perturbations that create the larger difference among 6 cell lines such as Belinostat (one of the positive controls), MNL 2238, and Oprozomib.  In the end, accuracy in predicting such chemicals may be important. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fa5aac54f6b18ab37d2cdaa3e3c4c921e%2Fsg_plot.png?generation=1701966669627694&amp;alt=media\" alt=\"\"></p>\n<h1>Model</h1>\n<h3>Simpler model as in the background section.</h3>\n<ol>\n<li><p>Solve linear system, transformer, from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the DE using the chemicals tested in the two cell lines (15 plus 2 positive control).</p>\n<p>The transformer can be computed by multiplying a pseudo-inverse of DE<em>_i_core</em> with DE<em>_j_core</em> from the left.</p>\n<p>DE<em>_i_core</em>  x T = DE<em>_j_core</em><br>\nT = DE<em>_i_core</em>-t * DE<em>_j_core</em></p>\n<p>where DE<em>_i_core</em>-t is a pseudo-inverse of DE<em>_i_core</em> </p></li>\n<li><p>Apply the transformer to the DE of the base cell line/target chemicals and get the DE of the target cell line/chemicals.</p></li>\n</ol>\n<p>The simpler model consumes a lot of memory &gt; 20Gb, and it cannot be performed on the free version of the Saturn Cloud, which I had used for convenience, I also propose an alternative solution using SVD with a projection space.</p>\n<h3>SVD projection</h3>\n<p>It first reduces the gene dimension (18,211) to the full rank dimension of the entire data set (614) by SVD, create the transformer in the projected space, and applies the transformer in projected space from known cell to target cell on the target chemicals, reconstruct the original gene dimension of the prediction.</p>\n<ol>\n<li>Project the DE data by SVD with the whole dimension.</li>\n<li>Solve linear system from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the projected DE using the chemicals tested in the two cell lines (15 plus 2 positive control).</li>\n<li>Apply the transformer to the projected DE of the base cell line/target chemicals and get the projected DE of the target cell line/chemicals.</li>\n<li>Inverse the projected DE of the target cell line/chemicals to the predicted DE.</li>\n</ol>\n<h1>Robustness, Code, and Reproducibility</h1>\n<p>Please refer to <a href=\"https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582\" target=\"_blank\">my notebook</a>: the code can run on the notebook, and the result is deterministic, so reproducible. </p>\n<h1>Key Findings from the results</h1>\n<p>This method indeed can be used to diagnose the similarity of the \"master rule\" among the cell lines for scientific insight, and my result suggests that NK cells, along with T cells CD4+ can be stronger predictor cell lines for B Cell and Myeloid cells.   Interestingly, T cells CD4+ is a better predictor for B cell while NK cell is a better predictor for Myeloid cells.</p>\n<p>This finding may align with the biological findings - <a href=\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3072878/\" target=\"_blank\">NK cells can be derived from the myeloid lineage, Blood - The Journal of the American Society of Hematology 2011 3548 </a> and <a href=\"https://www.cell.com/trends/immunology/fulltext/S1471-4906(21)00117-4\" target=\"_blank\">T follicular helper cells cognately guide differentiation of antigen primed B cells in secondary lymphoid tissues - Trends in immunology 42.8 2021</a>.</p>\n<h3>Prediction by single base cell line</h3>\n<table>\n<thead>\n<tr>\n<th>base cell</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>NK cells</td>\n<td>0.784</td>\n<td>0.596</td>\n</tr>\n<tr>\n<td>T cells CD4+</td>\n<td>0.775</td>\n<td>0.607</td>\n</tr>\n<tr>\n<td>T cells CD8+</td>\n<td>0.959</td>\n<td>0.706</td>\n</tr>\n<tr>\n<td>T regulatory cells</td>\n<td>0.834</td>\n<td>0.680</td>\n</tr>\n</tbody>\n</table>\n<h3>Prediction by two base cell lines</h3>\n<table>\n<thead>\n<tr>\n<th>base cell/target cell</th>\n<th>public</th>\n<th>private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Predict B Cell by NK, Myeloid cells by T4</td>\n<td>0.786</td>\n<td>0.616</td>\n</tr>\n<tr>\n<td>Predict B Cell by T4, Myeloid cells by NK</td>\n<td>0.773</td>\n<td>0.587</td>\n</tr>\n</tbody>\n</table>\n<ol>\n<li><p>Differences in DE among certain cell lines can be captured linearly very well, and the linear relation is well transferable across different chemical responses, 0.768 in private and 0.582 in public</p></li>\n<li><p>DE of NK cell and T cells CD4+ cells are quite predictable for that of B Cell and Myeloid cells</p></li>\n<li><p>T cells CD4+ is a better predictor for B cell, and NK cells is a better predictor for Myeloid cell</p></li>\n</ol>",
      "rawMarkdown": "# Intro\n\nI share a linear algebra method, an unbiased and reproducible approach, resulting in decent Private/Public scores of 0.768/0.582.  My final submission is an ensemble of this [Linear Algebra](https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582) approach along with AE NN mimicking the linear algebra approach by NN, whose joint weight is  0.70, and the combination of the public notebooks, [Pyboost](https://www.kaggle.com/code/makio323/pyboost-secret-grandmaster-s-tool-0-592), [NN](https://www.kaggle.com/code/makio323/fork-of-nlp-regression-12a31a-0-594), and [Linear SVR](https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain), with the total weight of 0.3.   It turns out that the pure Linear Algebra model gets the best private leaderboard score.\n\nIn the first half of this competition, I struggled to overcome the wall of a 0.600 public score.  Some lucky runs of some NN models got over the wall, but not always.  There was also the second formidable wall of 0.585, which blended the results of public models.\n\nAfter the first half, I came up with this linear algebra approach and overcame the walls; the prediction is deterministic and reproducible and helped me to move on. \n\nThe code of this linear model is available at [my code notebook](https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582). \n\n\n# Biological Hypothesis\n\nIt is a bit old, before the deep neural network, I had research experience using linear algebra to predict the missing values in a matrix -  [Missing Value Expectation of Matrix Data by Fixed Rank Approximation Algorithm](https://www.cs.uic.edu/~mtamura/MakioTamuraMasterProject.pdf), and it may inspire me.\n\nAn assumption behind this method is that the differential expressions (DEs) of 18,211 genes at one cell line (e.g. NK cells) can be linearly transferable to those of another cell line (e.g. B Cells) on the same chemical perturbation.\n\nA chemical perturbation triggers a complex activity interaction among the 18,211 genes, and different chemical perturbations make different activity patterns, resulting in various DEs from the same baseline condition.  However, there would be an unseen “master rule” to govern these interactions on each cell line.  If the master rule of one cell line (e.g. NK cells) could be similar to another cell (e.g. B cells), DEs on the same chemical perturbation could be predictable from one to another.  Even not knowing the master rule of each cell line, their relationship among cell lines could be captured such that\n\n- *f*(DE*_i_c*) = DE*_j_c*\n\nwhere DE*_i_x*and DE*_j_x* are the differential expressions of cell line *i* (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and *j* ('B cells', 'Myeloid cells') of chemical *c* perturbation, *f* is some special function.\n\n\nMy approach is to assume that the linear system can be the proxy function, and to solve a system may provide the \"transfer\" such that\n\n- DE*_i_core*  x T = DE*_j_core*\n\nwhere DE*_i_core* and DE*_j_core* are *m* x *n* matrix, *m* is the number of chemicals shared by cel line *i* (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and *j* ('B cells', 'Myeloid cells') as shown below (including positive controls) and n is the number of genes (18,211).  T is an *n* x *n* matrix, considered a transformer matrix from one cell line to another.\n\n\n  NK cells B cells 17\n  NK cells Myeloid cells 17\n  T cells CD4+ B cells 17\n  T cells CD4+ Myeloid cells 17\n  T cells CD8+ B cells 15\n  T cells CD8+ Myeloid cells 15\n  T regulatory cells B cells 17\n  T regulatory cells Myeloid cells 17\n\n\nOnce the transformer T is solved on the core chemicals, it may be applied to predict DEs of the target cell (e.g. B cells) from the known cell (e.g. NK cells)\n\n- Prediction of DE*_i_target* = DE*_j_target*  x T\n\nwhere DE*_i_target* is the DEs of prediction cell line i (B cells and Myeloid cells) of the target chemicals, 128 for B cells and 127 for Myoloid cells, and DE*_j_target* is the DEs of known cell line j (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) of the target chemicals.\n\nThis approach may provide a robust and unbiased prediction.\n\n\n\n# Observation\n\nUsing SVD to get the 1st and 2nd projected expression of the shared chemicals across 6 cell lines, and visualize the disperse among them (17 including positive controls, missing chemicals in T8 is replaced with those of T4).  Well, it is a bit difficult to recognize a clear pattern.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fb9f7d6b1595bae80513bbcd2fa786bb4%2Fs_plot.png?generation=1701966557440410&alt=media)\n\n\nHere is a grid plot of the previous one by chemicals.  In most of the chemicals, there is not much difference among 6 cell cline.  However, NK cell seems to be similar to the B Cell and Myoloid cell on the chemical perturbations that create the larger difference among 6 cell lines such as Belinostat (one of the positive controls), MNL 2238, and Oprozomib.  In the end, accuracy in predicting such chemicals may be important. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fa5aac54f6b18ab37d2cdaa3e3c4c921e%2Fsg_plot.png?generation=1701966669627694&alt=media)\n\n\n\n\n# Model\n\n### Simpler model as in the background section.\n\n1. Solve linear system, transformer, from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the DE using the chemicals tested in the two cell lines (15 plus 2 positive control).\n\n\tThe transformer can be computed by multiplying a pseudo-inverse of DE*_i_core* with DE*_j_core* from the left.\n\n\tDE*_i_core*  x T = DE*_j_core*\n\tT = DE*_i_core*-t * DE*_j_core*\n\n\twhere DE*_i_core*-t is a pseudo-inverse of DE*_i_core* \n\t\n2. Apply the transformer to the DE of the base cell line/target chemicals and get the DE of the target cell line/chemicals.\n\nThe simpler model consumes a lot of memory > 20Gb, and it cannot be performed on the free version of the Saturn Cloud, which I had used for convenience, I also propose an alternative solution using SVD with a projection space.\n\n\n\n\n### SVD projection\n\nIt first reduces the gene dimension (18,211) to the full rank dimension of the entire data set (614) by SVD, create the transformer in the projected space, and applies the transformer in projected space from known cell to target cell on the target chemicals, reconstruct the original gene dimension of the prediction.\n\n1. Project the DE data by SVD with the whole dimension.\n2. Solve linear system from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the projected DE using the chemicals tested in the two cell lines (15 plus 2 positive control).\n3. Apply the transformer to the projected DE of the base cell line/target chemicals and get the projected DE of the target cell line/chemicals.\n4. Inverse the projected DE of the target cell line/chemicals to the predicted DE.\n\n\n\n# Robustness, Code, and Reproducibility\n\n\nPlease refer to [my notebook](https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582): the code can run on the notebook, and the result is deterministic, so reproducible. \n\n\n# Key Findings from the results\n\nThis method indeed can be used to diagnose the similarity of the \"master rule\" among the cell lines for scientific insight, and my result suggests that NK cells, along with T cells CD4+ can be stronger predictor cell lines for B Cell and Myeloid cells.   Interestingly, T cells CD4+ is a better predictor for B cell while NK cell is a better predictor for Myeloid cells.\n\nThis finding may align with the biological findings - [NK cells can be derived from the myeloid lineage, Blood - The Journal of the American Society of Hematology 2011 3548 ](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3072878/) and [T follicular helper cells cognately guide differentiation of antigen primed B cells in secondary lymphoid tissues - Trends in immunology 42.8 2021](https://www.cell.com/trends/immunology/fulltext/S1471-4906(21)00117-4).\n\n\n\n### Prediction by single base cell line\n\n| base cell | public | private |\n| --- | --- | --- |\n|  NK cells | 0.784 | 0.596 |\n|  T cells CD4+ | 0.775 | 0.607 |\n|  T cells CD8+ | 0.959 | 0.706 |\n|  T regulatory cells | 0.834 | 0.680 |\n\n\n### Prediction by two base cell lines\n\n| base cell/target cell | public | private |\n| --- | --- | --- |\n|  Predict B Cell by NK, Myeloid cells by T4 | 0.786 | 0.616 |\n|  Predict B Cell by T4, Myeloid cells by NK | 0.773 | 0.587 |\n\n\n\n1. Differences in DE among certain cell lines can be captured linearly very well, and the linear relation is well transferable across different chemical responses, 0.768 in private and 0.582 in public\n\n2. DE of NK cell and T cells CD4+ cells are quite predictable for that of B Cell and Myeloid cells\n\n3. T cells CD4+ is a better predictor for B cell, and NK cells is a better predictor for Myeloid cell",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2552731": "# Intro\n\nI share a linear algebra method, an unbiased and reproducible approach, resulting in decent Private/Public scores of 0.768/0.582.  My final submission is an ensemble of this [Linear Algebra](https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582) approach along with AE NN mimicking the linear algebra approach by NN, whose joint weight is  0.70, and the combination of the public notebooks, [Pyboost](https://www.kaggle.com/code/makio323/pyboost-secret-grandmaster-s-tool-0-592), [NN](https://www.kaggle.com/code/makio323/fork-of-nlp-regression-12a31a-0-594), and [Linear SVR](https://www.kaggle.com/code/mehrankazeminia/1-op2-eda-linearsvr-regressorchain), with the total weight of 0.3.   It turns out that the pure Linear Algebra model gets the best private leaderboard score.\n\nIn the first half of this competition, I struggled to overcome the wall of a 0.600 public score.  Some lucky runs of some NN models got over the wall, but not always.  There was also the second formidable wall of 0.585, which blended the results of public models.\n\nAfter the first half, I came up with this linear algebra approach and overcame the walls; the prediction is deterministic and reproducible and helped me to move on. \n\nThe code of this linear model is available at [my code notebook](https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582). \n\n\n# Biological Hypothesis\n\nIt is a bit old, before the deep neural network, I had research experience using linear algebra to predict the missing values in a matrix -  [Missing Value Expectation of Matrix Data by Fixed Rank Approximation Algorithm](https://www.cs.uic.edu/~mtamura/MakioTamuraMasterProject.pdf), and it may inspire me.\n\nAn assumption behind this method is that the differential expressions (DEs) of 18,211 genes at one cell line (e.g. NK cells) can be linearly transferable to those of another cell line (e.g. B Cells) on the same chemical perturbation.\n\nA chemical perturbation triggers a complex activity interaction among the 18,211 genes, and different chemical perturbations make different activity patterns, resulting in various DEs from the same baseline condition.  However, there would be an unseen “master rule” to govern these interactions on each cell line.  If the master rule of one cell line (e.g. NK cells) could be similar to another cell (e.g. B cells), DEs on the same chemical perturbation could be predictable from one to another.  Even not knowing the master rule of each cell line, their relationship among cell lines could be captured such that\n\n- *f*(DE*_i_c*) = DE*_j_c*\n\nwhere DE*_i_x*and DE*_j_x* are the differential expressions of cell line *i* (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and *j* ('B cells', 'Myeloid cells') of chemical *c* perturbation, *f* is some special function.\n\n\nMy approach is to assume that the linear system can be the proxy function, and to solve a system may provide the \"transfer\" such that\n\n- DE*_i_core*  x T = DE*_j_core*\n\nwhere DE*_i_core* and DE*_j_core* are *m* x *n* matrix, *m* is the number of chemicals shared by cel line *i* (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) and *j* ('B cells', 'Myeloid cells') as shown below (including positive controls) and n is the number of genes (18,211).  T is an *n* x *n* matrix, considered a transformer matrix from one cell line to another.\n\n\n  NK cells B cells 17\n  NK cells Myeloid cells 17\n  T cells CD4+ B cells 17\n  T cells CD4+ Myeloid cells 17\n  T cells CD8+ B cells 15\n  T cells CD8+ Myeloid cells 15\n  T regulatory cells B cells 17\n  T regulatory cells Myeloid cells 17\n\n\nOnce the transformer T is solved on the core chemicals, it may be applied to predict DEs of the target cell (e.g. B cells) from the known cell (e.g. NK cells)\n\n- Prediction of DE*_i_target* = DE*_j_target*  x T\n\nwhere DE*_i_target* is the DEs of prediction cell line i (B cells and Myeloid cells) of the target chemicals, 128 for B cells and 127 for Myoloid cells, and DE*_j_target* is the DEs of known cell line j (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) of the target chemicals.\n\nThis approach may provide a robust and unbiased prediction.\n\n\n\n# Observation\n\nUsing SVD to get the 1st and 2nd projected expression of the shared chemicals across 6 cell lines, and visualize the disperse among them (17 including positive controls, missing chemicals in T8 is replaced with those of T4).  Well, it is a bit difficult to recognize a clear pattern.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fb9f7d6b1595bae80513bbcd2fa786bb4%2Fs_plot.png?generation=1701966557440410&alt=media)\n\n\nHere is a grid plot of the previous one by chemicals.  In most of the chemicals, there is not much difference among 6 cell cline.  However, NK cell seems to be similar to the B Cell and Myoloid cell on the chemical perturbations that create the larger difference among 6 cell lines such as Belinostat (one of the positive controls), MNL 2238, and Oprozomib.  In the end, accuracy in predicting such chemicals may be important. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F577034%2Fa5aac54f6b18ab37d2cdaa3e3c4c921e%2Fsg_plot.png?generation=1701966669627694&alt=media)\n\n\n\n\n# Model\n\n### Simpler model as in the background section.\n\n1. Solve linear system, transformer, from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the DE using the chemicals tested in the two cell lines (15 plus 2 positive control).\n\n\tThe transformer can be computed by multiplying a pseudo-inverse of DE*_i_core* with DE*_j_core* from the left.\n\n\tDE*_i_core*  x T = DE*_j_core*\n\tT = DE*_i_core*-t * DE*_j_core*\n\n\twhere DE*_i_core*-t is a pseudo-inverse of DE*_i_core* \n\t\n2. Apply the transformer to the DE of the base cell line/target chemicals and get the DE of the target cell line/chemicals.\n\nThe simpler model consumes a lot of memory > 20Gb, and it cannot be performed on the free version of the Saturn Cloud, which I had used for convenience, I also propose an alternative solution using SVD with a projection space.\n\n\n\n\n### SVD projection\n\nIt first reduces the gene dimension (18,211) to the full rank dimension of the entire data set (614) by SVD, create the transformer in the projected space, and applies the transformer in projected space from known cell to target cell on the target chemicals, reconstruct the original gene dimension of the prediction.\n\n1. Project the DE data by SVD with the whole dimension.\n2. Solve linear system from a base cell line (NK cells, T cells CD4+, T cells CD8+, T regulatory cells) to a target cell line ('B cells', 'Myeloid cells') of the projected DE using the chemicals tested in the two cell lines (15 plus 2 positive control).\n3. Apply the transformer to the projected DE of the base cell line/target chemicals and get the projected DE of the target cell line/chemicals.\n4. Inverse the projected DE of the target cell line/chemicals to the predicted DE.\n\n\n\n# Robustness, Code, and Reproducibility\n\n\nPlease refer to [my notebook](https://www.kaggle.com/code/makio323/24th-using-linear-algebra-priv-pub-0-768-0-582): the code can run on the notebook, and the result is deterministic, so reproducible. \n\n\n# Key Findings from the results\n\nThis method indeed can be used to diagnose the similarity of the \"master rule\" among the cell lines for scientific insight, and my result suggests that NK cells, along with T cells CD4+ can be stronger predictor cell lines for B Cell and Myeloid cells.   Interestingly, T cells CD4+ is a better predictor for B cell while NK cell is a better predictor for Myeloid cells.\n\nThis finding may align with the biological findings - [NK cells can be derived from the myeloid lineage, Blood - The Journal of the American Society of Hematology 2011 3548 ](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3072878/) and [T follicular helper cells cognately guide differentiation of antigen primed B cells in secondary lymphoid tissues - Trends in immunology 42.8 2021](https://www.cell.com/trends/immunology/fulltext/S1471-4906(21)00117-4).\n\n\n\n### Prediction by single base cell line\n\n| base cell | public | private |\n| --- | --- | --- |\n|  NK cells | 0.784 | 0.596 |\n|  T cells CD4+ | 0.775 | 0.607 |\n|  T cells CD8+ | 0.959 | 0.706 |\n|  T regulatory cells | 0.834 | 0.680 |\n\n\n### Prediction by two base cell lines\n\n| base cell/target cell | public | private |\n| --- | --- | --- |\n|  Predict B Cell by NK, Myeloid cells by T4 | 0.786 | 0.616 |\n|  Predict B Cell by T4, Myeloid cells by NK | 0.773 | 0.587 |\n\n\n\n1. Differences in DE among certain cell lines can be captured linearly very well, and the linear relation is well transferable across different chemical responses, 0.768 in private and 0.582 in public\n\n2. DE of NK cell and T cells CD4+ cells are quite predictable for that of B Cell and Myeloid cells\n\n3. T cells CD4+ is a better predictor for B cell, and NK cells is a better predictor for Myeloid cell"
  }
}