{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nCreate models on different subsamples and compare features importances.\n\nLasso models\n\nWe fix alpha and create/study Lasso models for all targets. \nAlpha = 1 seems to be over strong for all targets, so it can give some look on the most strongest genes.\nFor alpha = 1 all targets - 4.0 minutes. For 106 targets all coefs are zeros. For alpha = 1 - correlation between target std and predictions characteristics are very strong: \n0.880610\t0.604685\t0.581836\t0.848555\t0.849982\t0.979509\t0.977644\t0.790846\t\n\n\n### Versions:\n\n1-4 - alpha = 1, train_size = 0.1 - very strong regularization, not many targets \"survive\", quite fast - 3 minutes for all \n\n5 train_size = 0.5 (changed), alpha = 1 (same) \n\n6 train_size = 0.9 (changed), alpha = 1 (same) \n\n7 alpha = 0.1 (changed), train_size = 0.1 (changed) \n\n8 train_size = 0.5 (changed) , alpha = 0.1 (same), \n\n9 train_size = 0.9 (changed) , alpha = 0.1 (same), \n\n10 train_size = 0.1 (changed) , alpha = 0.01 (changed), \n\n","metadata":{}},{"cell_type":"markdown","source":"# Key params","metadata":{}},{"cell_type":"code","source":"alpha = 0.01\ntrain_size = 0.1\nstr_method = 'Lasso '\n#n_trials = 100 #fast_mode = 0\nimport time\nt0start = time.time() ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Some Findings \n\n### Often see RPS4Y1 - as predictor - it is gender related: Y-chromosome, almost same similar see: ENSG00000128040_SPINK2 - probably also gender related \n\n###  These CD have more than 5 inflential RNA for alpha = 1 \n    Here \"Corr\" is correlation of the prediction to target \n    \n     11.0 CD48  Corr: 0.67 ['CD48', 'NKG7', 'CEBPD', 'LY86', 'SUCNR1', 'GPR183', 'RAB27A', 'S100A11', 'RAB31', 'FAM107B']\n    9.0 CD33  Corr: 0.46 ['RNASE2', 'RPS23', 'SRGN', 'FAM107B', 'CALR', 'S100A11', 'LGALS1', 'RPS4Y1', 'TPT1']\n    7.0 CD44  Corr: 0.51 ['SPINK2', 'BST2', 'NPDC1', 'CD44', 'GATA1', 'IFITM3', 'HBD']\n    17.0 CD31  Corr: 0.6 ['SPINK2', 'RPS4Y1', 'KCNH2', 'C1QTNF4', 'NPDC1', 'CNRIP1', 'HOPX', 'SMIM24', 'XIST', 'HLA-A']\n    16.0 CD32  Corr: 0.75 ['SPINK2', 'AKAP12', 'FCGR2B', 'LMNA', 'FCGR2A', 'GSTP1', 'RPL5', 'GATA1', 'FCER1A', 'PMP22']\n    7.0 CD62L Corr: 0.73 ['SPINK2', 'SELL', 'C1QTNF4', 'GATA1', 'GLIPR1', 'BST2', 'LMNA']\n    13.0 CD11a Corr: 0.6 ['C1QTNF4', 'GLIPR1', 'CD48', 'IGHM', 'NKG7', 'LSP1', 'RAB31', 'SPINK2', 'SPNS3', 'ITGAL']\n    6.0 CD244 Corr: 0.56 ['SPINK2', 'XIST', 'RPS4Y1', 'SMIM24', 'AVP', 'HBD']\n    15.0 CD41  Corr: 0.83 ['CMTM5', 'THBS1', 'HBD', 'BST2', 'ARHGAP6', 'LMNA', 'ITGB3', 'RAB6B', 'CTTNBP2', 'GP9']\n    12.0 CD49b Corr: 0.65 ['IFITM3', 'IGHM', 'RPS4X', 'SPINK2', 'TM4SF1', 'IFI16', 'HOPX', 'TPT1', 'PRDX1', 'FAM30A']\n    6.0 CD38  Corr: 0.58 ['CD38', 'RNASE2', 'CEBPE', 'RNASE3', 'CLEC12A', 'FAM30A']\n    10.0 CD36  Corr: 0.76 ['CD36', 'ANK1', 'HBD', 'HBA1', 'HBBP1', 'LMNA', 'BST2', 'SMIM1', 'SPTA1', 'RHAG']\n\n### KLF1, GATA1 clearly appears for some CDs, note that top correlated are:  ['CD71' 'CD115' 'CD88'] : \n\n    5.0 CD71  Corr: 0.67 ['KLF1', 'TPT1', 'B2M', 'GATA1', 'YBX1']\n\n    4.0 CD82  Corr: 0.47 ['S100A6', 'GATA1', 'ITGA2B', 'FCER1A']\n\n    2.0 CD88  Corr: 0.64 ['KLF1', 'GATA1']\n    \n    \n### MAFB  - 1.0 CD11c Corr: 0.13 ['MAFB'] - transcription factor\n\n    Transcription factor MafB also known as V-maf musculoaponeurotic fibrosarcoma oncogene homolog B is a protein that in humans is encoded by the MAFB gene. This gene maps to chromosome 20q11.2-q13.1, consists of a single exon and spans around 3 kb.[5][6]\n    MafB is a basic leucine zipper (bZIP) transcription factor that plays an important role in the regulation of lineage-specific hematopoiesis. The encoded nuclear protein represses ETS1-mediated transcription of erythroid-specific genes in myeloid cells.[6]\n    Recently, single-nucleotide polymorphisms (SNPs) near MAFB have been found associated with nonsyndromic cleft lip and palate. [9] The GENEVA Cleft Consortium study, a genomewide association study involving 1,908 case-parent trios from Europe, the United States, China, Taiwan, Singapore, Korea, and the Philippines, first identified MAFB as being associated with cleft lip and/or palate with stronger genome-wide significance in Asian than European populations.\n\n### TPT1 - the only predictor for CD54 for alpha = 1: CD54  Corr: 0.41 ['TPT1']\n\n    Translationally Controlled Tumor Protein (TCTP) is a protein that in humans is encoded by the TPT1 gene.[4][5][6] The TPT1 gene is mapped 13q12-q1413 in the Chromosome 13.[5] The human gene contains five introns and six exons, The TPT1-gene contains a promoter with a canonical TATA-box and several promoter elements, which are well-conserved in mammals.[7] The assay with reporter gene exhibits a strong promoter activity comparable to viral promoters.[8]\n\n    TCTP protein is also referred to as Q23,[9] P21,[10] P23,[11] histamine releasing factor (HRF),[12] and fortilin.[13] TCTP is a multifunctional and highly conserved protein that existed ubiquitously in different eukaryote species and distributed widely in various tissues and cell types.[14]\n\n    Human translationally controlled tumor protein (hTCTP) is a growth-related, calcium-binding protein.[15]\n","metadata":{}},{"cell_type":"markdown","source":"# From Previous notebooks\n\n\n#### Findings:\n\n**alpha = 0.1** for CD36 does not depend on subsample size ( train_size - parameter ) \n\n**Score: 0.79-0.8** correlation (Pearson) with \"y_true\" (CD36) for all subsample sizes \n\n**Lasso seems not always works fine**  For target CD44 alpha = 0.01 is optimal (from the list), but the number of non-zero features is almost 10_000 - that seems quite inappropriately large. Enforcing alpha = 0.1 a bit decreases the quality non-zero features becomes - about 400. Intersection of 100 trials - 147. \n\n\n#### Versions: \n\n    1,2 - n_trials = 10\n    3 - n_trials = 100 \n    4 - draft save  \n    5 - n_trials = 10, train_size = 0.75\n    6 - n_trials = 100, train_size = 0.75\n    7 - n_trials = 20, train_size = 0.75    \n    8 - n_trials = 100, train_size = 0.5 - repeat   \n    9 - n_trials = 10, train_size = 0.5 - repeat   \n    10,11 - n_trials = 10 , train_size = 0.1 \n    12 - n_trials = 100 , train_size = 0.1 \n    13 - n_trials = 100 , train_size = 0.9\n    14,15,16,17 - n_trials = 10 , train_size = 0.9\n    17,18 - crash\n    19 - n_trials = 100 , train_size = 0.9 - 9 hours \n    20,21,22  - n_trials = 10,50,100, train_size = 0.9 - repeat\n    23,24,25 - crash\n    26,27,28  - n_trials = 10 train_size = 0.95 \n    29  quick save - 2 trials, fast mode , train_size = 0.5 - 7 minutes \n    30  - n_trials = 100 train_size = 0.95  # Cancelled after 12 hours\n    31  - n_trials = 100 train_size = 0.9   # Ran in 9 hours and 28 minutes\n    32  - n_trials = 100 train_size = 0.75  # Ran in 8 hours and 34 minutes\n    33  - n_trials = 100 train_size = 0.5   # Ran in 3 hours and 06 minutes\n    34  - n_trials = 100 train_size = 0.1   # Ran in 3 hours and 49 minutes\n    \n    35,36  - n_trials = 10 train_size = 0.1 - 30.4 minutes , 860 non-zero features - for 1 trial, but intersection between two - 126 and later falls down to 18: [823, 126, 59, 30, 21, 20, 19, 19, 19, 18]\n    37 - n_trials = 10 train_size = 0.5\n    38 - n_trials = 10 train_size = 0.75\n    39 - n_trials = 10 train_size = 0.9\n    40 - n_trials = 10 train_size = 0.95 - 120 minutes, 0.802 0.807 158 non-zero feautures  Intersection: [158, 150, 144, 138, 135, 133, 133, 131, 129, 129], Correlations - 0.99 \n    \n    \n    ====================================================================\n    CD44\n    ====================================================================\n    41 n_trials = 10 train_size = 0.1 # Ran in 36 minutes and 27 seconds # 0.675972\t0.744553\t580.500000  alpha:  0.1\n    42 n_trials = 10 train_size = 0.5 # Ran in 34 minutes and 42 seconds # 0.683030\t0.693261\t409.700000  alpha = 0.1\n\n    43 n_trials = 100 train_size = 0.1 # Ran in 2 hours and 56 minutes\n    44 n_trials = 100 train_size = 0.5 # Ran in 6 hours and 22 minutes\n    45 n_trials = 100 train_size = 0.75 # Cancelled after 12 hours\n    46 n_trials = 100 train_size = 0.9 # Cancelled after 12 hours\n    47 n_trials = 100 train_size = 0.95 # Cancelled after 12 hours\n\n    48 n_trials = 10 train_size = 0.75 # Ran in 6 hours and 14 minutes #  0.693954\t0.804693\t9104.300000 alpha: 0.01 (! different alpha ! , standard alpha = 0.1 gives a bit less score: 0.679396 with 391 features )  [9151, 5776, 4465, 3709, 3269, 2928, 2684, 2459, 2286, 2152]\n    49 n_trials = 10 train_size = 0.9  # Ran in 5 hours and 10 minutes\n    50 n_trials = 10 train_size = 0.95 # Run in 4.75 hours # 0.703858\t0.786153\t8613.700000 # alpha: 0.01 \n                All: [8555, 7174, 6502, 6102, 5821, 5619, 5438, 5308, 5176, 5074]\n                Count Non-zero median importances  8762\n                Average count genes non-zero for trials:  8613.7\n                Count genes non-zero for all trials:  5074\n                number of non-zero coeffients common for Pairs of trials:   7330.29\n                number of non-zero coeffients common for TRIPLES of trials:   6531.725\n\n    \n    ====================================================================\n    CD36, CD44, \n    ====================================================================\n    51  - CD36, n_trials = 100 train_size = 0.1 # 3.21 hours # 0.790715\t0.867522\t864.480000 # \n                Count Non-zero median importances  84\n                Average count genes non-zero for trials:  864.48\n                First 30: [897, 119, 49, 30, 27, 26, 22, 20, 18, 16, 15, 15, 14, 13, 11, 9, 9, 9, 9, 8, 8, 8, 7, 7, 7, 7, 7, 7, 7, 7]\n                Last 30:  [5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5]\n                \n    52  - CD44  n_trials = 20 train_size = 0.9  # Expexted is about 10 hours - compare to V49 # 7.52 hours\n                # 0.700470\t0.790130\t8719.250000\n                Count Non-zero median importances  8514\n                Average count genes non-zero for trials:  8719.25\n                Count genes non-zero for all trials:  3246\n                All: [8717, 6662, 5787, 5242, 4876, 4607, 4399, 4229, 4091, 3974, 3870, 3770, 3691, 3601, 3511, 3442, 3384, 3332, 3285, 3246]\n                \n    53  - CD44  n_trials = 15 train_size = 0.75  # Expexted is about 10 hours - compare to V48 # 7.58 hours\n        All: [9018, 5717, 4382, 3657, 3160, 2854, 2619, 2407, 2251, 2120, 2033, 1938, 1858, 1770, 1714]\n        Count Non-zero median importances  7920\n        Average count genes non-zero for trials:  9110\n        Count genes non-zero for all trials:  1714\n\n         Alpha           Train      Test         N_nonzero_feature\n        1.000000e-03\t0.849821\t0.611386\t19942.0\n        1.000000e-02\t0.803358\t0.698067\t9261.0\n        1.000000e-01\t0.688624\t0.683882\t407.0\n        1.000000e+00\t0.498581\t0.499372\t6.0    \n\n    ====================================================================\n    CD44, enforce alpha = 0.1 \n    ====================================================================\n    \n    54 - CD44  enforce alpha = 0.1 , n_trials = 100 train_size = 0.5 # 6.17 hours # \t0.682333\t0.693402\t412.080000\n            Count Non-zero median importances  306\n            Average count genes non-zero for trials:  412.08\n            Count genes non-zero for all trials:  69\n            [390, 251, 202, 174, 163, 154, 148, 139, 131, 129, 127, 124, 122, 121, 118, 117, 115, 111, 107, 105, 104, 104, 102, 100, 98, 96, 96, 96, 94, 94, 92, 89, 88, 87, 86, 86, 86, 85, 84, 84, 84, 84, 84, 84, 84, 84, 84, 83, 82, 82, 82, 81, 81, 81, 80, 79, 79, 79, 78, 78, 78, 78, 77, 77, 77, 77, 76, 76, 76, 76, 76, 76, 76, 75, 75, 75, 75, 75, 74, 74, 73, 72, 72, 72, 72, 72, 72, 72, 71, 71, 71, 71, 71, 71, 71, 70, 70, 70, 70, 69]\n\n    55 - CD44  enforce alpha = 0.1 , n_trials = 100 train_size = 0.75  # 7.28 hours # 0.682817\t0.689567\t400.210000\n        Count Non-zero median importances  363\n        Average count genes non-zero for trials:  400.21\n        Count genes non-zero for all trials:  147\n            [406, 305, 271, 250, 233, 221, 217, 212, 203, 198, 193, 190, 190, 187, 185, 184, 183, 181, 180, 178, 177, 176, 176, 176, 175, 173, 172, 171, 171, 171, 171, 170, 170, 168, 166, 165, 165, 164, 164, 163, 163, 163, 163, 163, 161, 161, 161, 160, 160, 160, 160, 160, 160, 159, 158, 156, 156, 156, 156, 156, 155, 155, 154, 153, 153, 153, 153, 152, 152, 152, 152, 151, 151, 151, 151, 151, 150, 150, 150, 149, 149, 148, 148, 148, 148, 147, 147, 147, 147, 147, 147, 147, 147, 147, 147, 147, 147, 147, 147, 147]\n    \n    56 - CD44  enforce alpha = 0.1 , n_trials = 100 train_size = 0.9 # 8 hours, 24 minutes  \n        # 0.681424\t0.688602\t395.720000\n        Count Non-zero median importances  378\n        Average count genes non-zero for trials:  395.72\n        Count genes non-zero for all trials:  216\n        All: [388, 328, 315, 304, 291, 286, 284, 275, 271, 267, 265, 264, 263, 261, 259, 258, 256, 250, 250, 250, 249, 248, 248, 247, 247, 245, 243, 242, 242, 240, 240, 240, 238, 238, 238, 237, 237, 236, 235, 235, 235, 235, 234, 234, 232, 232, 232, 232, 232, 232, 232, 230, 230, 230, 230, 229, 228, 228, 227, 226, 226, 226, 226, 226, 226, 225, 225, 225, 225, 225, 225, 225, 224, 224, 224, 223, 223, 223, 222, 222, 221, 221, 221, 221, 221, 221, 221, 221, 221, 220, 220, 218, 217, 217, 217, 217, 217, 216, 216, 216]\n        \n    57  - CD44  enforce alpha = 0.1 , n_trials = 100 train_size = 0.1  \n        \n    ====================================================================\n    CD88 \n    ====================================================================\n        \n    58  - CD88  enforce alpha = 0.1 , n_trials = 100 train_size = 0.1  \n    59  - CD88  enforce alpha = 0.1 , n_trials = 100 train_size = 0.9 \n    \n    60  - CD88  NO enforce alpha , n_trials = 100 train_size = 0.1  \n    61  - CD88  NO enforce alpha , n_trials = 100 train_size = 0.5  \n    \n    \n    \n# From the previous notebook with Ridge model: \nimport numpy as np\nimport pandas as pd \n\nprint('Unfortunately intersection of lists from 10 trials and 100 both 47 gives 37, so not fully coincident.  ')\n\n\n## From Ridge CD36\n\n### From Version 3: \n    l_10trials_top100_gives47intersection = ['ENSG00000135218_CD36', 'ENSG00000229988_HBBP1', 'ENSG00000112077_RHAG', 'ENSG00000137801_THBS1', 'ENSG00000130303_BST2', 'ENSG00000109099_PMP22', 'ENSG00000110092_CCND1', 'ENSG00000029534_ANK1', 'ENSG00000223609_HBD', 'ENSG00000129824_RPS4Y1', 'ENSG00000204103_MAFB', 'ENSG00000101162_TUBB1', 'ENSG00000235169_SMIM1', 'ENSG00000267279_AC090409.1', 'ENSG00000164946_FREM1', 'ENSG00000168685_IL7R', 'ENSG00000244734_HBB', 'ENSG00000073464_CLCN4', 'ENSG00000197993_KEL', 'ENSG00000113924_HGD', 'ENSG00000185198_PRSS57', 'ENSG00000139174_PRICKLE1', 'ENSG00000166091_CMTM5', 'ENSG00000047648_ARHGAP6', 'ENSG00000196565_HBG2', 'ENSG00000065534_MYLK', 'ENSG00000165682_CLEC1B', 'ENSG00000107984_DKK1', 'ENSG00000088053_GP6', 'ENSG00000206172_HBA1', 'ENSG00000250361_GYPB', 'ENSG00000170873_MTSS1', 'ENSG00000187609_EXD3', 'ENSG00000107130_NCS1', 'ENSG00000115461_IGFBP5', 'ENSG00000169704_GP9', 'ENSG00000103522_IL21R', 'ENSG00000138135_CH25H', 'ENSG00000072274_TFRC', 'ENSG00000169403_PTAFR', 'ENSG00000047597_XK', 'ENSG00000073737_DHRS9', 'ENSG00000077984_CST7', 'ENSG00000251002_AC244502.1', 'ENSG00000092621_PHGDH', 'ENSG00000169071_ROR2', 'ENSG00000213719_CLIC1']\n\n#### From Version 2: \nl_100trials_top300_gives47intersection = ['ENSG00000135218_CD36', 'ENSG00000229988_HBBP1', 'ENSG00000130303_BST2', 'ENSG00000112077_RHAG', 'ENSG00000109099_PMP22', 'ENSG00000137801_THBS1', 'ENSG00000110092_CCND1', 'ENSG00000029534_ANK1', 'ENSG00000223609_HBD', 'ENSG00000164946_FREM1', 'ENSG00000073464_CLCN4', 'ENSG00000235169_SMIM1', 'ENSG00000267279_AC090409.1', 'ENSG00000101162_TUBB1', 'ENSG00000129824_RPS4Y1', 'ENSG00000168685_IL7R', 'ENSG00000026025_VIM', 'ENSG00000197993_KEL', 'ENSG00000244734_HBB', 'ENSG00000113924_HGD', 'ENSG00000166091_CMTM5', 'ENSG00000047648_ARHGAP6', 'ENSG00000139174_PRICKLE1', 'ENSG00000185198_PRSS57', 'ENSG00000065534_MYLK', 'ENSG00000115461_IGFBP5', 'ENSG00000169704_GP9', 'ENSG00000088053_GP6', 'ENSG00000187609_EXD3', 'ENSG00000175084_DES', 'ENSG00000133742_CA1', 'ENSG00000107130_NCS1', 'ENSG00000072274_TFRC', 'ENSG00000047597_XK', 'ENSG00000131981_LGALS3', 'ENSG00000205542_TMSB4X', 'ENSG00000170873_MTSS1', 'ENSG00000198673_FAM19A2', 'ENSG00000169403_PTAFR', 'ENSG00000198400_NTRK1', 'ENSG00000251002_AC244502.1', 'ENSG00000092621_PHGDH', 'ENSG00000213719_CLIC1', 'ENSG00000169071_ROR2', 'ENSG00000251562_MALAT1', 'ENSG00000091409_ITGA6', 'ENSG00000174788_PCP2']\nm = pd.Series(index = l_10trials_top100_gives47intersection, dtype = float).index.isin(l_100trials_top300_gives47intersection )\ns = list( pd.Series(index = l_10trials_top100_gives47intersection, dtype = float)[m].index )\nprint(s[:2])\nprint( len(s), len(l_10trials_top100_gives47intersection), len(l_100trials_top300_gives47intersection ) )\nl = [t.split('_')[1] for t in s ]\nprint('Intersection:', len(l), l)\n\n    \n    ","metadata":{}},{"cell_type":"markdown","source":"# Preparations","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-13T10:43:42.276100Z","iopub.execute_input":"2023-01-13T10:43:42.277031Z","iopub.status.idle":"2023-01-13T10:43:42.298131Z","shell.execute_reply.started":"2023-01-13T10:43:42.276987Z","shell.execute_reply":"2023-01-13T10:43:42.296584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:43:42.300289Z","iopub.execute_input":"2023-01-13T10:43:42.301239Z","iopub.status.idle":"2023-01-13T10:43:43.004414Z","shell.execute_reply.started":"2023-01-13T10:43:42.301191Z","shell.execute_reply":"2023-01-13T10:43:43.002561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"%%time\nfilename_rna_data = '/kaggle/input/open-problems-multimodal/train_cite_inputs.h5'\ndf_rna = pd.read_hdf(filename_rna_data)\ndisplay(df_rna) \n\n#%%time\ndf_y = pd.read_hdf('/kaggle/input/open-problems-multimodal/train_cite_targets.h5')\ndisplay(df_y)\n\nfn = '/kaggle/input/open-problems-multimodal/metadata.csv'\ndf_meta = pd.read_csv(fn, index_col = 0 )\ndf_meta\n# Cut only train cite-seq part: \nd = pd.DataFrame(index = df_y.index)\nprint(d.shape)\ndf_meta = d.join(df_meta, how = 'left')\ndf_meta","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:43:43.006706Z","iopub.execute_input":"2023-01-13T10:43:43.007132Z","iopub.status.idle":"2023-01-13T10:44:43.964220Z","shell.execute_reply.started":"2023-01-13T10:43:43.007096Z","shell.execute_reply":"2023-01-13T10:44:43.962953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print( dict( df_meta.value_counts('cell_type') ))","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:44:43.966079Z","iopub.execute_input":"2023-01-13T10:44:43.966593Z","iopub.status.idle":"2023-01-13T10:44:43.982759Z","shell.execute_reply.started":"2023-01-13T10:44:43.966548Z","shell.execute_reply":"2023-01-13T10:44:43.980524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Modeling ","metadata":{}},{"cell_type":"code","source":"# y = df_y[target_name].values\n# print(y.shape, type(y) )\n\nX = ( df_rna.values )\nprint(X.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:44:43.984516Z","iopub.execute_input":"2023-01-13T10:44:43.985209Z","iopub.status.idle":"2023-01-13T10:44:43.993102Z","shell.execute_reply.started":"2023-01-13T10:44:43.985165Z","shell.execute_reply":"2023-01-13T10:44:43.991234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.preprocessing import StandardScaler\nscaler = StandardScaler()\n\nX = scaler.fit_transform( X )\nprint(X.shape)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:44:43.995912Z","iopub.execute_input":"2023-01-13T10:44:43.996599Z","iopub.status.idle":"2023-01-13T10:45:08.764588Z","shell.execute_reply.started":"2023-01-13T10:44:43.996540Z","shell.execute_reply":"2023-01-13T10:45:08.762974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.linear_model import LassoCV\nfrom sklearn.linear_model import Lasso\nfrom sklearn.linear_model import RidgeCV\nfrom sklearn.linear_model import Ridge\n\nfrom sklearn.model_selection import cross_val_predict\nfrom sklearn.model_selection import KFold \nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:45:08.767703Z","iopub.execute_input":"2023-01-13T10:45:08.768413Z","iopub.status.idle":"2023-01-13T10:45:08.933924Z","shell.execute_reply.started":"2023-01-13T10:45:08.768370Z","shell.execute_reply":"2023-01-13T10:45:08.932394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train model","metadata":{}},{"cell_type":"code","source":"%%time \nn_cd_to_consider = 140\n\nlen_y = len(df_y)\np = np.random.permutation( len_y )\nN = int(len_y *  train_size )\nIX_train = np.arange(len_y)[p][:N]\nIX_test =  np.arange(len_y)[p][N:]\nprint(\" len(IX_train), len(IX_test):\",  len(IX_train), len(IX_test) )\n\n\n\nt0 = time.time()\n\n\nmodel = Lasso(alpha = alpha )\n\nprint(model)\n\nmodel.fit(X[IX_train,:], df_y.values[IX_train,:n_cd_to_consider ])\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:48:30.856372Z","iopub.execute_input":"2023-01-13T10:48:30.856831Z","iopub.status.idle":"2023-01-13T10:50:13.460592Z","shell.execute_reply.started":"2023-01-13T10:48:30.856797Z","shell.execute_reply":"2023-01-13T10:50:13.459055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time \ny_pred = model.predict(X )\n# df_y.values[IX_train,:2]\n    ","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:10.143496Z","iopub.execute_input":"2023-01-13T10:54:10.144105Z","iopub.status.idle":"2023-01-13T10:54:16.552572Z","shell.execute_reply.started":"2023-01-13T10:54:10.144065Z","shell.execute_reply":"2023-01-13T10:54:16.550435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred.shape","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:16.556329Z","iopub.execute_input":"2023-01-13T10:54:16.557600Z","iopub.status.idle":"2023-01-13T10:54:16.567539Z","shell.execute_reply.started":"2023-01-13T10:54:16.557531Z","shell.execute_reply":"2023-01-13T10:54:16.566062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom sklearn.metrics import r2_score\nfrom sklearn.metrics import mean_squared_error\n\ndf_models = pd.DataFrame()\nfor i in range(y_pred.shape[1]):\n    #col = df_y.columns[i]\n    col = df_y.columns[i] + ' Lasso alpha'+str(alpha) + ' TrainSize'+str(int(train_size*100) ) +'%'  \n\n    df_models.loc[col, 'n_nonZeros'] =  (model.coef_[i,:] !=0).sum()\n\n    v1 = df_y.values[IX_test, i ]\n    v2 = y_pred[IX_test,i]\n    c = np.corrcoef(v1,v2)[0,1]\n    df_models.loc[col, 'Corr Pearson Test'] = c\n    print(c, col, 'r2:', np.round( r2_score(v1,v2) , 3 ), 'Test')\n\n    v1 = df_y.values[IX_train, i ]\n    v2 = y_pred[IX_train,i]\n    c = np.corrcoef(v1,v2)[0,1]\n    df_models.loc[col, 'Corr Pearson Train'] = c\n    print('  ', c, col, 'r2:', np.round( r2_score(v1,v2) , 3 ), 'Train')\n\n    \n    v1 = df_y.values[IX_test, i ];     v2 = y_pred[IX_test,i]\n    df_models.loc[col, 'r2 Test'] = r2_score(v1,v2)\n    v1 = df_y.values[IX_train, i ];     v2 = y_pred[IX_train,i]\n    df_models.loc[col, 'r2 Train'] = r2_score(v1,v2)\n    \n    v1 = df_y.values[IX_test, i ];     v2 = y_pred[IX_test,i]\n    df_models.loc[col, 'RMSE Test'] = mean_squared_error(v1,v2, squared = False)\n    v1 = df_y.values[IX_train, i ];     v2 = y_pred[IX_train,i]\n    df_models.loc[col, 'RMSE Train'] = mean_squared_error(v1,v2, squared = False)\n\n    df_models.loc[col, 'Mean Target'] = df_y[df_y.columns[i]].mean()\n    df_models.loc[col, 'Std Target'] = df_y[df_y.columns[i]].std()\n    \n    df_models.loc[col, 'Target'] = df_y.columns[i]\n    df_models.loc[col, 'Alpha'] = alpha\n    df_models.loc[col, 'Train Size'] = train_size\n    \n    \n    \nprint();print('Top 15')\ndisplay( df_models.sort_values('Corr Pearson Test', ascending = False).head(15) )\nprint();print('Tail 15')\ndisplay( df_models.sort_values('Corr Pearson Test', ascending = False).tail(15) )\n","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:16.570058Z","iopub.execute_input":"2023-01-13T10:54:16.571187Z","iopub.status.idle":"2023-01-13T10:54:19.318266Z","shell.execute_reply.started":"2023-01-13T10:54:16.571109Z","shell.execute_reply":"2023-01-13T10:54:19.316846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn = 'model_stat_NIPS22_'+str_method.replace(' ','')+ '_alpha'+str(alpha).replace('.','d')+'_TrainSize'+str(int(train_size*100) ) +'percent'+'.csv'\ndf_models.to_csv( fn )\nprint(fn)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:19.322149Z","iopub.execute_input":"2023-01-13T10:54:19.322628Z","iopub.status.idle":"2023-01-13T10:54:19.336802Z","shell.execute_reply.started":"2023-01-13T10:54:19.322589Z","shell.execute_reply":"2023-01-13T10:54:19.335675Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Look on models statistics","metadata":{}},{"cell_type":"code","source":"df_models.head(3)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:19.338215Z","iopub.execute_input":"2023-01-13T10:54:19.338627Z","iopub.status.idle":"2023-01-13T10:54:19.366176Z","shell.execute_reply.started":"2023-01-13T10:54:19.338591Z","shell.execute_reply":"2023-01-13T10:54:19.364902Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m = df_models['Corr Pearson Test'].isnull().sum()\nprint(m.sum())","metadata":{"execution":{"iopub.status.busy":"2023-01-13T11:13:41.059078Z","iopub.execute_input":"2023-01-13T11:13:41.059505Z","iopub.status.idle":"2023-01-13T11:13:41.068761Z","shell.execute_reply.started":"2023-01-13T11:13:41.059471Z","shell.execute_reply":"2023-01-13T11:13:41.067320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models.describe()","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:19.368430Z","iopub.execute_input":"2023-01-13T10:54:19.368823Z","iopub.status.idle":"2023-01-13T10:54:19.426856Z","shell.execute_reply.started":"2023-01-13T10:54:19.368789Z","shell.execute_reply":"2023-01-13T10:54:19.425515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col = 'Corr Pearson Test'\nfig = plt.figure(figsize = (20,6))\nplt.plot(df_models[col].values, '*-' , label = col)\nplt.legend(fontsize = 20)\nplt.grid()\nplt.show()\n\n\nfor col in ['n_nonZeros','r2 Test', 'RMSE Test', 'Std Target', 'Mean Target' ]:\n    fig = plt.figure(figsize = (20,6))\n    plt.plot(df_models[col].values, '*-' , label = col)\n    plt.legend(fontsize = 20)\n    plt.grid()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:19.430833Z","iopub.execute_input":"2023-01-13T10:54:19.431625Z","iopub.status.idle":"2023-01-13T10:54:21.033854Z","shell.execute_reply.started":"2023-01-13T10:54:19.431582Z","shell.execute_reply":"2023-01-13T10:54:21.032373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m = df_models['Corr Pearson Test'].isnull().sum()\nprint(m.sum())","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:21.035531Z","iopub.execute_input":"2023-01-13T10:54:21.036365Z","iopub.status.idle":"2023-01-13T10:54:21.044100Z","shell.execute_reply.started":"2023-01-13T10:54:21.036315Z","shell.execute_reply":"2023-01-13T10:54:21.042219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models.sort_values('Corr Pearson Test', ascending = False).head(20)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:21.046641Z","iopub.execute_input":"2023-01-13T10:54:21.047242Z","iopub.status.idle":"2023-01-13T10:54:21.085553Z","shell.execute_reply.started":"2023-01-13T10:54:21.047187Z","shell.execute_reply":"2023-01-13T10:54:21.084481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models.sort_values('Corr Pearson Test', ascending = False).tail(20)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:21.090114Z","iopub.execute_input":"2023-01-13T10:54:21.091205Z","iopub.status.idle":"2023-01-13T10:54:21.124980Z","shell.execute_reply.started":"2023-01-13T10:54:21.091147Z","shell.execute_reply":"2023-01-13T10:54:21.123367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models.corr()","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:21.126487Z","iopub.execute_input":"2023-01-13T10:54:21.127002Z","iopub.status.idle":"2023-01-13T10:54:21.154660Z","shell.execute_reply.started":"2023-01-13T10:54:21.126950Z","shell.execute_reply":"2023-01-13T10:54:21.153027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Print Save Importances ","metadata":{}},{"cell_type":"code","source":"df_importances = pd.DataFrame(index = df_rna.columns )\n\nfor i in range( model.coef_.shape[0] ):\n    col = df_y.columns[i] + ' Lasso alpha'+str(alpha) + ' TrainSize'+str(int(train_size*100) ) +'%'  \n    df_importances[col] = model.coef_[i,:]\n    #print(col, (df_importances[col]!=0).sum())\n    l = list( df_importances[col].sort_values(ascending =False, key = abs).index )[:15]\n    l = [t.split('_')[1] for t in l]\n    #print( l )\n    #print()\ndf_importances","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:21.156494Z","iopub.execute_input":"2023-01-13T10:54:21.156966Z","iopub.status.idle":"2023-01-13T10:54:22.011669Z","shell.execute_reply.started":"2023-01-13T10:54:21.156918Z","shell.execute_reply":"2023-01-13T10:54:22.010214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn = 'importances_NIPS22_'+str_method.replace(' ','')+ '_alpha'+str(alpha).replace('.','d')+'_TrainSize'+str(int(train_size*100) ) +'percent'+'.csv'\ndf_importances.to_csv( fn )\nprint(fn)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:22.013113Z","iopub.execute_input":"2023-01-13T10:54:22.013506Z","iopub.status.idle":"2023-01-13T10:54:23.735187Z","shell.execute_reply.started":"2023-01-13T10:54:22.013473Z","shell.execute_reply":"2023-01-13T10:54:23.733781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_rows', 500)\npd.set_option('display.max_columns', 500)\npd.set_option('display.width', 1000)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:23.736491Z","iopub.execute_input":"2023-01-13T10:54:23.736837Z","iopub.status.idle":"2023-01-13T10:54:23.745054Z","shell.execute_reply.started":"2023-01-13T10:54:23.736806Z","shell.execute_reply":"2023-01-13T10:54:23.743235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_models.T","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:23.746840Z","iopub.execute_input":"2023-01-13T10:54:23.747366Z","iopub.status.idle":"2023-01-13T10:54:23.888944Z","shell.execute_reply.started":"2023-01-13T10:54:23.747312Z","shell.execute_reply":"2023-01-13T10:54:23.887403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d = pd.DataFrame()\nfor i,col in enumerate(df_importances.columns):\n    l = list( df_importances.sort_values(col,ascending=False, key = abs).index)\n    d[col] = [t.split('_')[1] for t in l] \n#d2 = pd.concat( (df_models[ ['n_nonZeros']].T , d ) , axis = 0 )\nd2 = pd.concat( (df_models.T , d ) , axis = 0 )\n    \nfn = 'genesSortedByImpotance_NIPS22_'+str_method.replace(' ','')+ '_alpha'+str(alpha).replace('.','d')+'_TrainSize'+str(int(train_size*100) ) +'percent'+'.csv'\nd.to_csv( fn )\nprint(fn)\nfn = 'genesSortedImpotanceAndInformation_NIPS22_'+str_method.replace(' ','')+ '_alpha'+str(alpha).replace('.','d')+'_TrainSize'+str(int(train_size*100) ) +'percent'+'.csv'\nd2.to_csv( fn )\nprint(fn)\n\nd2.head(60) ","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:23.890860Z","iopub.execute_input":"2023-01-13T10:54:23.892114Z","iopub.status.idle":"2023-01-13T10:54:31.406114Z","shell.execute_reply.started":"2023-01-13T10:54:23.892057Z","shell.execute_reply":"2023-01-13T10:54:31.404285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"m = d2.iloc[0,:] != 0\nprint(m.sum())\nd2.loc[:,m].head(60)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:31.408463Z","iopub.execute_input":"2023-01-13T10:54:31.409589Z","iopub.status.idle":"2023-01-13T10:54:31.675085Z","shell.execute_reply.started":"2023-01-13T10:54:31.409530Z","shell.execute_reply":"2023-01-13T10:54:31.673570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(d.shape[1]):\n    col = d.columns[i]\n    n = d2.iloc[0,i]\n    t = np.round( d2.loc['Corr Pearson Test',:].iat[i] , 2)\n    if n>0:\n        kk = int( np.min([n, 30] ) )\n        print( n, col[:5],'Corr:', t,  list(d[col].iloc[:kk]) )\n        print()","metadata":{"execution":{"iopub.status.busy":"2023-01-13T11:08:17.005854Z","iopub.execute_input":"2023-01-13T11:08:17.007009Z","iopub.status.idle":"2023-01-13T11:08:17.057100Z","shell.execute_reply.started":"2023-01-13T11:08:17.006960Z","shell.execute_reply":"2023-01-13T11:08:17.055796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(d.shape[1]):\n    col = d.columns[i]\n    n = d2.iloc[0,i]\n    t = np.round( d2.loc['Corr Pearson Test',:].iat[i] , 2)\n    if n>5:\n        kk = int( np.min([n, 10] ) )\n        print( n, col[:5],'Corr:', t,  list(d[col].iloc[:kk]) )","metadata":{"execution":{"iopub.status.busy":"2023-01-13T11:06:49.779752Z","iopub.execute_input":"2023-01-13T11:06:49.780187Z","iopub.status.idle":"2023-01-13T11:06:49.823401Z","shell.execute_reply.started":"2023-01-13T11:06:49.780153Z","shell.execute_reply":"2023-01-13T11:06:49.821540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Look in importances","metadata":{}},{"cell_type":"markdown","source":"## What genes are most frequenly important among all targets ","metadata":{}},{"cell_type":"code","source":"v = (df_importances != 0).sum(axis=1).sort_values(ascending = False)\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v.values[:100], '*-' , label = 'Frequencies for top important genes')\nplt.legend(fontsize = 20)\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:31.677301Z","iopub.execute_input":"2023-01-13T10:54:31.678243Z","iopub.status.idle":"2023-01-13T10:54:31.946138Z","shell.execute_reply.started":"2023-01-13T10:54:31.678197Z","shell.execute_reply":"2023-01-13T10:54:31.944844Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(df_importances != 0).sum(axis=1).sort_values(ascending = False).head(50)","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:31.947740Z","iopub.execute_input":"2023-01-13T10:54:31.948632Z","iopub.status.idle":"2023-01-13T10:54:31.978711Z","shell.execute_reply.started":"2023-01-13T10:54:31.948579Z","shell.execute_reply":"2023-01-13T10:54:31.976868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tt = time.time() - t0start\nprint('%.1f seconds ( = %.1f minutes, = %.2f hours) passed'%( tt, tt/60, tt/3600 ) )","metadata":{"execution":{"iopub.status.busy":"2023-01-13T10:54:31.980375Z","iopub.execute_input":"2023-01-13T10:54:31.980794Z","iopub.status.idle":"2023-01-13T10:54:31.988326Z","shell.execute_reply.started":"2023-01-13T10:54:31.980758Z","shell.execute_reply":"2023-01-13T10:54:31.986926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}