{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is about ?\n\nCompare feature importances obtained by different methods.\n\n\n### Versions\n\n\n#### 1-6\n\n    compare lists obtained by permutation importance method from sklearn \n    https://scikit-learn.org/stable/modules/permutation_importance.html\n    Applied to Ridge model for CD36 protein for all 22K features (rna) - dataset Kaggle/NIPS22 Competition on multiomodal single cell data: https://www.kaggle.com/competitions/open-problems-multimodal\n    The computations were made by Pavel Chekanov (https://www.kaggle.com/hydrophonyx), they took more 17 hours each. \n    \n    The idea is - region where parabolic approximation works well - is complete disconcordance between two orderings of importances  - that means these features are probably NOT important since two methods disagree on them as much as possible in theory. And the region in the begining - where linear approximation works - idicates features which selected by both methods.\n    So we can estimate number of truly important features. Roughly speaking taking as boundary point - the point where parabolic approximation is better than linear.  \n    \n    For theoretical justifications see:\n    \nhttps://mathoverflow.net/q/437545/10446    https://mathoverflow.net/q/437569/10446\n    \n    Parabolic approximation begin to work well starting from n = 200 \n    So we can conclude that only about 200 features might be important. That is less than other esimates, but might be indication that Ridge+Permutation importance are not suitable for such task. \n    \n    Original recipe to choose important features via sklearn-permuation importance - to choose ones with positive importances. It appears to be that method is not that good in that example. For both lists the number of positives is about 10k from 22k - which is at least 5 times more than many other estimates. And interesection of the two leasts gives about 6.8k features which indicates that methods are not well agreed, confirming impression that \"Ridge+Permutation importance are not suitable for such task\"\n    \n#### 7 \n    Added comparison with ordering by correlation. Roughly speaking 20% goes to intersection  \n    \n    Also with Spearman correlation and mutual information. \n    We see that some genes can be catched by model importance, but not by these univariable importances,\n    that is quite interesting effect. ","metadata":{}},{"cell_type":"markdown","source":"# Preparations and data loading","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\n\nl = [] \ni = 0\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    l = l + [ dirname ] \n    for filename in filenames:\n        i += 1\n        if i<=10:\n            print('One of the first 10 files: ', os.path.join(dirname, filename))\nprint()\nprint('Dirs:', set(l))        \nprint('\\n n_files:', i)\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-14T20:33:38.629470Z","iopub.execute_input":"2023-01-14T20:33:38.630265Z","iopub.status.idle":"2023-01-14T20:33:38.648314Z","shell.execute_reply.started":"2023-01-14T20:33:38.630223Z","shell.execute_reply":"2023-01-14T20:33:38.646933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:15:02.813949Z","iopub.execute_input":"2023-01-14T19:15:02.814540Z","iopub.status.idle":"2023-01-14T19:15:03.580979Z","shell.execute_reply.started":"2023-01-14T19:15:02.814488Z","shell.execute_reply":"2023-01-14T19:15:03.580047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn1 = '/kaggle/input/cd36-sklearn-perm-importance/CD36_permutation_imp_ridge_v2.csv' # Here alpha = 1e6 - better than alpha = 1\nfn2 = '/kaggle/input/cd36permutationimportance/CD36_permutation_imp_ridge.csv' # Here alpha = 1\ndf1 = pd.read_csv(fn1, index_col = 0)\ndf1","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:15:03.582284Z","iopub.execute_input":"2023-01-14T19:15:03.583423Z","iopub.status.idle":"2023-01-14T19:15:03.681477Z","shell.execute_reply.started":"2023-01-14T19:15:03.583382Z","shell.execute_reply":"2023-01-14T19:15:03.680307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2 = pd.read_csv(fn2, index_col = 0)\ndf2","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:15:03.683976Z","iopub.execute_input":"2023-01-14T19:15:03.684352Z","iopub.status.idle":"2023-01-14T19:15:03.758423Z","shell.execute_reply.started":"2023-01-14T19:15:03.684322Z","shell.execute_reply":"2023-01-14T19:15:03.757040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = df1.sort_values(df1.columns[1], ascending = False)\ndf2 = df2.sort_values(df1.columns[1], ascending = False)\ndisplay(df1.head(5))\ndisplay(df2.head(5))\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:15:03.760730Z","iopub.execute_input":"2023-01-14T19:15:03.761143Z","iopub.status.idle":"2023-01-14T19:15:03.791941Z","shell.execute_reply.started":"2023-01-14T19:15:03.761108Z","shell.execute_reply":"2023-01-14T19:15:03.790640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis of importances concordances ","metadata":{}},{"cell_type":"code","source":"k = 100\ns = set(df1.index[:k] ) & set(df2.index[:k] ) \nlen(s)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:22:45.735453Z","iopub.execute_input":"2023-01-14T19:22:45.735915Z","iopub.status.idle":"2023-01-14T19:22:45.745379Z","shell.execute_reply.started":"2023-01-14T19:22:45.735869Z","shell.execute_reply":"2023-01-14T19:22:45.744210Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for k in [10,25,50, 75, 100, 1000, 2000, 4000]:\n    s = set(df1.index[:k] ) & set(df2.index[:k] ) \n    len(s)\n    print(k, len(s))","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:22:45.747995Z","iopub.execute_input":"2023-01-14T19:22:45.748987Z","iopub.status.idle":"2023-01-14T19:22:45.759128Z","shell.execute_reply.started":"2023-01-14T19:22:45.748943Z","shell.execute_reply":"2023-01-14T19:22:45.757926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nl = []\nfor k in range(df1.shape[0]): # [10,25,50, 75, 100, 1000, 2000, 4000]:\n    s = set(df1.index[:k] ) & set(df2.index[:k] ) \n    l.append(len(s) )\n    # print(k, len(s))","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:22:45.761497Z","iopub.execute_input":"2023-01-14T19:22:45.761991Z","iopub.status.idle":"2023-01-14T19:24:16.343458Z","shell.execute_reply.started":"2023-01-14T19:22:45.761946Z","shell.execute_reply":"2023-01-14T19:24:16.342107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize = (20,6))\nplt.plot(l ,label = 'original data - counts intersection' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x,y,2)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'parabolic approximation' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x[:50],y[:50],1)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'linear approximation  based on 0-50 points' )\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\n\nfig = plt.figure(figsize = (20,10))\nplt.plot(l ,label = 'original data - counts intersection' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x,y,2)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'parabolic approximation' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x[:50],y[:50],1)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'linear approximation based on 0-50 points' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x[:150],y[:150],1)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'linear approximation  based on 0-150 points' )\n\nplt.xlim([0,500])\nplt.ylim([0,250])\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:16.346080Z","iopub.execute_input":"2023-01-14T19:24:16.346564Z","iopub.status.idle":"2023-01-14T19:24:16.999340Z","shell.execute_reply.started":"2023-01-14T19:24:16.346520Z","shell.execute_reply":"2023-01-14T19:24:16.998062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize = (20,10))\nplt.plot(l ,label = 'original data - counts intersection' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x,y,2)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'parabolic approximation' )\n\ny = np.array(l)\nx = np.arange(len(y))\np = np.polyfit(x[:250],y[:250],1)\nprint(p)\nplt.plot(np.polyval(p,x),label = 'linear approximation  based on 0-250 points' )\n\nplt.xlim([0,3000])\nplt.ylim([0,1500])\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:17.000902Z","iopub.execute_input":"2023-01-14T19:24:17.001270Z","iopub.status.idle":"2023-01-14T19:24:17.312096Z","shell.execute_reply.started":"2023-01-14T19:24:17.001238Z","shell.execute_reply":"2023-01-14T19:24:17.310827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analysis of input importances ","metadata":{}},{"cell_type":"code","source":"display( df1.describe() )\ndisplay( df2.describe() )\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:17.315940Z","iopub.execute_input":"2023-01-14T19:24:17.316289Z","iopub.status.idle":"2023-01-14T19:24:17.343180Z","shell.execute_reply.started":"2023-01-14T19:24:17.316259Z","shell.execute_reply":"2023-01-14T19:24:17.342252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Original method to choose important features as those which are give positve number value - gives about 10k features - which by all estimates is unreasonable 5-10 times. And intersection of those chosen gives 6878 which also indicates that 10k is unreasonable","metadata":{}},{"cell_type":"code","source":"df = df1\ncol = df.columns[1]; v = df[col];\nprint('Positive', 'Zero', 'Negative')\nprint((v>0).sum(), (v == 0).sum(), (v<0).sum() )\ns1 = set( v[( v>0)].index )\n\ndf = df2\ncol = df.columns[1]; v = df[col];\nprint((v>0).sum(), (v == 0).sum(), (v<0).sum() )\ns2 = set( v[( v>0)].index )\n\ns = s1 & s2 \nprint('Interesection of selected by 1 and 2 methods: ', len(s) )","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:17.344471Z","iopub.execute_input":"2023-01-14T19:24:17.345023Z","iopub.status.idle":"2023-01-14T19:24:17.363461Z","shell.execute_reply.started":"2023-01-14T19:24:17.344990Z","shell.execute_reply":"2023-01-14T19:24:17.362620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = df1\ncol = df.columns[1]; v1 = df[col];\ndf = df2\ncol = df.columns[1]; v2 = df[col];\n\n(v1>0)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:17.364738Z","iopub.execute_input":"2023-01-14T19:24:17.365250Z","iopub.status.idle":"2023-01-14T19:24:17.379736Z","shell.execute_reply.started":"2023-01-14T19:24:17.365218Z","shell.execute_reply":"2023-01-14T19:24:17.378406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Direct Compare two importances plots","metadata":{}},{"cell_type":"code","source":"\nv1 = df1[df1.columns[1] ]\nv2 = df2[df2.columns[1] ]\n\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v1.values, '*-' ,label = 'feuture importances 1' )\nplt.plot(v2.values, '*-' ,label = 'feuture importances 2' )\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v1.values, '*-' ,label = 'feuture importances 1' )\nplt.plot(v2.values, '*-' ,label = 'feuture importances 2' )\n\nplt.xlim([0,100])\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\nfig = plt.figure(figsize = (20,6))\nm1 = v1>0\nm2 = v2>0\n\nplt.plot(np.log(v1[m1].values[:1000]), '*-' ,label = 'Log feuture importances 1' )\nplt.plot(np.log(v2[m2].values[:1000]), '*-' ,label = 'Log feuture importances 2' )\n\n#plt.xlim([0,100])\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:17.380989Z","iopub.execute_input":"2023-01-14T19:24:17.381322Z","iopub.status.idle":"2023-01-14T19:24:18.163217Z","shell.execute_reply.started":"2023-01-14T19:24:17.381292Z","shell.execute_reply":"2023-01-14T19:24:18.161927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col1 = df1.columns[1]\nv = df1[col1]\nm = v > 0\nprint(m.sum(), (v == 0).sum(), (v<0).sum() )\n\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v.values, '*-' ,label = 'feuture importances' )\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v.values, '*-' ,label = 'feuture importances' )\n\nplt.xlim([0,100])\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\nfig = plt.figure(figsize = (20,6))\nm = v>0\nplt.plot(np.log(v[m].values[:1000]), '*-' ,label = 'Log feuture importances' )\n\n#plt.xlim([0,100])\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:18.164970Z","iopub.execute_input":"2023-01-14T19:24:18.165475Z","iopub.status.idle":"2023-01-14T19:24:19.103857Z","shell.execute_reply.started":"2023-01-14T19:24:18.165441Z","shell.execute_reply":"2023-01-14T19:24:19.102850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col2 = df2.columns[1]\nv = df2[col2]\nm = v > 0\nprint(m.sum(), (v == 0).sum(), (v<0).sum() )\n\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v.values, '*-' ,label = 'feuture importances' )\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\nfig = plt.figure(figsize = (20,6))\nplt.plot(v.values, '*-' ,label = 'feuture importances' )\n\nplt.xlim([0,100])\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n\nfig = plt.figure(figsize = (20,6))\nm = v>0\nplt.plot(np.log(v[m].values[:1000]), '*-' ,label = 'Log feuture importances' )\n\n#plt.xlim([0,100])\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:19.105004Z","iopub.execute_input":"2023-01-14T19:24:19.105826Z","iopub.status.idle":"2023-01-14T19:24:19.881000Z","shell.execute_reply.started":"2023-01-14T19:24:19.105792Z","shell.execute_reply":"2023-01-14T19:24:19.879921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show some important genes ","metadata":{}},{"cell_type":"code","source":"s= set(df1.index[:300]) & set(df2.index[:300]  )\nlen(s)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:24:34.617175Z","iopub.execute_input":"2023-01-14T19:24:34.617588Z","iopub.status.idle":"2023-01-14T19:24:34.625962Z","shell.execute_reply.started":"2023-01-14T19:24:34.617556Z","shell.execute_reply":"2023-01-14T19:24:34.624807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_rows', 500)\npd.set_option('display.max_columns', 500)\npd.set_option('display.width', 1000)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:25:48.248816Z","iopub.execute_input":"2023-01-14T19:25:48.250791Z","iopub.status.idle":"2023-01-14T19:25:48.257006Z","shell.execute_reply.started":"2023-01-14T19:25:48.250729Z","shell.execute_reply":"2023-01-14T19:25:48.255641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Show top 100 in intersection ","metadata":{}},{"cell_type":"code","source":"m = df1.index.isin(s)\ndf1[m]","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:25:49.178958Z","iopub.execute_input":"2023-01-14T19:25:49.179672Z","iopub.status.idle":"2023-01-14T19:25:49.200423Z","shell.execute_reply.started":"2023-01-14T19:25:49.179619Z","shell.execute_reply":"2023-01-14T19:25:49.199256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn3 = '/kaggle/input/nips22-cdrna-correlations/CD36_pearson_rna_all.csv'\ndf3 = pd.read_csv(fn3,index_col = 0)\ndf3 = df3.sort_values(df3.columns[0], key = abs, ascending = False )\ndf3","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:38:04.946345Z","iopub.execute_input":"2023-01-14T19:38:04.946781Z","iopub.status.idle":"2023-01-14T19:38:04.997755Z","shell.execute_reply.started":"2023-01-14T19:38:04.946747Z","shell.execute_reply":"2023-01-14T19:38:04.996192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn4 = '/kaggle/input/nips22-cdrna-correlations/CD36_spearman_rna_all.csv'\ndf4 = pd.read_csv(fn4,index_col = 0)\ndf4 = df4.sort_values(df4.columns[0], key = abs, ascending = False )\ndf4","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:36:09.371194Z","iopub.execute_input":"2023-01-14T19:36:09.371598Z","iopub.status.idle":"2023-01-14T19:36:09.423684Z","shell.execute_reply.started":"2023-01-14T19:36:09.371566Z","shell.execute_reply":"2023-01-14T19:36:09.422267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = set( df3.index[:100] ) & set( df4.index[:100])\nlen(s)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:38:17.518098Z","iopub.execute_input":"2023-01-14T19:38:17.518509Z","iopub.status.idle":"2023-01-14T19:38:17.526428Z","shell.execute_reply.started":"2023-01-14T19:38:17.518475Z","shell.execute_reply":"2023-01-14T19:38:17.525251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for k in [10,25,50, 75, 100, 1000, 2000, 4000]:\n    s = set(df3.index[:k] ) & set(df4.index[:k] ) \n    len(s)\n    print(k, len(s))","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:38:56.745388Z","iopub.execute_input":"2023-01-14T19:38:56.745818Z","iopub.status.idle":"2023-01-14T19:38:56.759369Z","shell.execute_reply.started":"2023-01-14T19:38:56.745784Z","shell.execute_reply":"2023-01-14T19:38:56.757891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = set( df1.set_index('0').index[:100] ) & set( df4.index[:100])\nprint( len(s) )\nprint('Intersection:')\n[t.split('_')[1] for t in df4.index[:100] if t in s ]\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:44:19.181337Z","iopub.execute_input":"2023-01-14T19:44:19.181784Z","iopub.status.idle":"2023-01-14T19:44:19.194375Z","shell.execute_reply.started":"2023-01-14T19:44:19.181749Z","shell.execute_reply":"2023-01-14T19:44:19.193354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# seems about interesection is about 20% of the lists \nfor k in [10,25,50, 75, 100, 1000, 2000, 4000]:\n    s = set(df1['0'].iloc[:k] ) & set(df4.index[:k] ) \n    len(s)\n    print(k, len(s))","metadata":{"execution":{"iopub.status.busy":"2023-01-14T19:41:44.921871Z","iopub.execute_input":"2023-01-14T19:41:44.922270Z","iopub.status.idle":"2023-01-14T19:41:44.933680Z","shell.execute_reply.started":"2023-01-14T19:41:44.922240Z","shell.execute_reply":"2023-01-14T19:41:44.932002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df3['Rank'] = (-df3['ABS_CD36']).rank()\ndf3","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:09:06.369323Z","iopub.execute_input":"2023-01-14T20:09:06.369711Z","iopub.status.idle":"2023-01-14T20:09:06.387458Z","shell.execute_reply.started":"2023-01-14T20:09:06.369680Z","shell.execute_reply":"2023-01-14T20:09:06.386334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d = df1.set_index('0')\nd = d.join(df3)\nd","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:09:07.134709Z","iopub.execute_input":"2023-01-14T20:09:07.135138Z","iopub.status.idle":"2023-01-14T20:09:07.164644Z","shell.execute_reply.started":"2023-01-14T20:09:07.135107Z","shell.execute_reply":"2023-01-14T20:09:07.163855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d.columns = ['Permutation Importance Ridge', 'Correlation', 'Abs Correlation', 'Rank Abs Correlation']\nd.tail(20)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:09:08.523775Z","iopub.execute_input":"2023-01-14T20:09:08.524171Z","iopub.status.idle":"2023-01-14T20:09:08.540608Z","shell.execute_reply.started":"2023-01-14T20:09:08.524140Z","shell.execute_reply":"2023-01-14T20:09:08.539502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize= (20,8))\ncol = 'Rank Abs Correlation'\nplt.plot(d[col].values[:30],'*-', label = col)\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:09:09.405165Z","iopub.execute_input":"2023-01-14T20:09:09.405827Z","iopub.status.idle":"2023-01-14T20:09:09.647962Z","shell.execute_reply.started":"2023-01-14T20:09:09.405793Z","shell.execute_reply":"2023-01-14T20:09:09.646743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df4['Rank CD36 Spearman'] = (-df4['ABS_CD36']).rank()\ndf4.head(1)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:09:10.879698Z","iopub.execute_input":"2023-01-14T20:09:10.880090Z","iopub.status.idle":"2023-01-14T20:09:10.893403Z","shell.execute_reply.started":"2023-01-14T20:09:10.880057Z","shell.execute_reply":"2023-01-14T20:09:10.892334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d2 = d.join(df4)\nd2 = d2.rename(columns={\"CD36\": \"CD36 Spearman\", \"ABS_CD36\": \"ABS CD36 Spearman\"})\nd2","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:09:19.360439Z","iopub.execute_input":"2023-01-14T20:09:19.360824Z","iopub.status.idle":"2023-01-14T20:09:19.388317Z","shell.execute_reply.started":"2023-01-14T20:09:19.360793Z","shell.execute_reply":"2023-01-14T20:09:19.387256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d2.iloc[20:30,:]","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:11:07.067512Z","iopub.execute_input":"2023-01-14T20:11:07.067934Z","iopub.status.idle":"2023-01-14T20:11:07.085848Z","shell.execute_reply.started":"2023-01-14T20:11:07.067900Z","shell.execute_reply":"2023-01-14T20:11:07.084730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize= (20,8))\ncol = 'Rank CD36 Spearman'\nplt.plot(d2[col].values[:150],'*-', label = col)\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:19:26.801305Z","iopub.execute_input":"2023-01-14T20:19:26.801822Z","iopub.status.idle":"2023-01-14T20:19:27.058195Z","shell.execute_reply.started":"2023-01-14T20:19:26.801774Z","shell.execute_reply":"2023-01-14T20:19:27.056884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compare with mutual information univariable importances ","metadata":{}},{"cell_type":"code","source":"# One of the first 10 files:  /kaggle/input/nips22-mutual-information/22050x140_round001.csv\n# One of the first 10 files:  /kaggle/input/nips22-mutual-information/22050x140_round01.csv\n# One of the first 10 files:  /kaggle/input/nips22-mutual-information/22050x140_round01_mutual_CD.csv\n# One of the first 10 files:  /kaggle/input/nips22-mutual-information/22050x140_round001_mutual_CD.csv\nfn5 = '/kaggle/input/nips22-mutual-information/22050x140_round01.csv'\ndf5 = pd.read_csv(fn5,index_col = 0)\ndf5\nfn6 = '/kaggle/input/nips22-mutual-information/22050x140_round001.csv'\ndf6 = pd.read_csv(fn6,index_col = 0)\ndf6","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:44:12.602110Z","iopub.execute_input":"2023-01-14T20:44:12.602548Z","iopub.status.idle":"2023-01-14T20:44:17.868927Z","shell.execute_reply.started":"2023-01-14T20:44:12.602510Z","shell.execute_reply":"2023-01-14T20:44:17.867727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"t = df5[['CD36'] ].copy()\nt['Rank CD36 MutualInfo'] = range(len(t))\nt = t.set_index('CD36')#, inplace = True)\nt\n\nt2 = df6[['CD36'] ].copy()\nt2['Rank CD36 MutualInfo-001'] = range(len(t))\nt2 = t2.set_index('CD36')#, inplace = True)\nt2\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:45:06.020024Z","iopub.execute_input":"2023-01-14T20:45:06.020475Z","iopub.status.idle":"2023-01-14T20:45:06.046607Z","shell.execute_reply.started":"2023-01-14T20:45:06.020438Z","shell.execute_reply":"2023-01-14T20:45:06.045445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d3 = d2.join(t)\nd3 = d3.join(t2)\nd3","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:45:21.071300Z","iopub.execute_input":"2023-01-14T20:45:21.071746Z","iopub.status.idle":"2023-01-14T20:45:21.117066Z","shell.execute_reply.started":"2023-01-14T20:45:21.071711Z","shell.execute_reply":"2023-01-14T20:45:21.115905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d3.iloc[20:30]","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:45:27.051638Z","iopub.execute_input":"2023-01-14T20:45:27.052556Z","iopub.status.idle":"2023-01-14T20:45:27.074745Z","shell.execute_reply.started":"2023-01-14T20:45:27.052496Z","shell.execute_reply":"2023-01-14T20:45:27.073582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize= (20,8))\ncol = 'Rank CD36 MutualInfo'#'Rank CD36 Spearman'\nplt.plot(d3[col].values[:150],'*-', label = col)\ncol = 'Rank CD36 MutualInfo-001'#'Rank CD36 Spearman'\nplt.plot(d3[col].values[:150],'*-', label = col)\n\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:47:17.986962Z","iopub.execute_input":"2023-01-14T20:47:17.987391Z","iopub.status.idle":"2023-01-14T20:47:18.301452Z","shell.execute_reply.started":"2023-01-14T20:47:17.987357Z","shell.execute_reply":"2023-01-14T20:47:18.300154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d3.columns","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:53:50.353072Z","iopub.execute_input":"2023-01-14T20:53:50.353723Z","iopub.status.idle":"2023-01-14T20:53:50.365480Z","shell.execute_reply.started":"2023-01-14T20:53:50.353672Z","shell.execute_reply":"2023-01-14T20:53:50.363794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d3[['Rank Abs Correlation','Rank CD36 Spearman', 'Rank CD36 MutualInfo',  'Rank CD36 MutualInfo-001']].min()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:56:29.137396Z","iopub.execute_input":"2023-01-14T20:56:29.137802Z","iopub.status.idle":"2023-01-14T20:56:29.152127Z","shell.execute_reply.started":"2023-01-14T20:56:29.137772Z","shell.execute_reply":"2023-01-14T20:56:29.151098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d3['Min Univar Rank'] = d3[['Rank Abs Correlation','Rank CD36 Spearman', 'Rank CD36 MutualInfo',  'Rank CD36 MutualInfo-001']].min(axis = 1)\nd3","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:56:48.675634Z","iopub.execute_input":"2023-01-14T20:56:48.676717Z","iopub.status.idle":"2023-01-14T20:56:48.708229Z","shell.execute_reply.started":"2023-01-14T20:56:48.676675Z","shell.execute_reply":"2023-01-14T20:56:48.706891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Candidates for genes not seen by \"univariable\" dependence searching methods. ENSG00000091409_ITGA6 , etc\n\nOnly seen by models ","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize= (20,8))\n# col = 'Rank CD36 MutualInfo'#'Rank CD36 Spearman'\n# plt.plot(d3[col].values[:150],'*-', label = col)\ncol = 'Min Univar Rank'# , 'Rank CD36 MutualInfo-001'#'Rank CD36 Spearman'\nplt.plot(d3[col].values[:150],'*-', label = col)\n\nplt.legend(fontsize = 20 )\nplt.grid()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:56:50.322953Z","iopub.execute_input":"2023-01-14T20:56:50.323722Z","iopub.status.idle":"2023-01-14T20:56:50.617889Z","shell.execute_reply.started":"2023-01-14T20:56:50.323646Z","shell.execute_reply":"2023-01-14T20:56:50.616897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col = 'Min Univar Rank'\nm = d3[col] > 3000\nd3.iloc[:150,:][m]","metadata":{"execution":{"iopub.status.busy":"2023-01-14T21:07:48.327169Z","iopub.execute_input":"2023-01-14T21:07:48.327673Z","iopub.status.idle":"2023-01-14T21:07:48.360492Z","shell.execute_reply.started":"2023-01-14T21:07:48.327618Z","shell.execute_reply":"2023-01-14T21:07:48.359171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"l = list( d3.iloc[:150,:][m].index )\nprint(len(l), l )","metadata":{"execution":{"iopub.status.busy":"2023-01-14T21:09:37.924721Z","iopub.execute_input":"2023-01-14T21:09:37.925179Z","iopub.status.idle":"2023-01-14T21:09:37.934187Z","shell.execute_reply.started":"2023-01-14T21:09:37.925144Z","shell.execute_reply":"2023-01-14T21:09:37.932772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"d3.head(20)","metadata":{"execution":{"iopub.status.busy":"2023-01-14T20:57:36.027292Z","iopub.execute_input":"2023-01-14T20:57:36.027744Z","iopub.status.idle":"2023-01-14T20:57:36.056396Z","shell.execute_reply.started":"2023-01-14T20:57:36.027710Z","shell.execute_reply":"2023-01-14T20:57:36.054967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}