{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**What are you trying to do in this notebook?**\n\nTo predict how DNA, RNA, and protein measurements co-vary in single cells as bone marrow stem cells develop into more mature blood cells. I will develop a model trained on a subset of 300,000-cell time course dataset of CD34+ hematopoietic stem and progenitor cells (HSPC) from four human donors at five time points generated for this competition by Cellarity, a cell-centric drug creation company.\n\nMy work will help accelerate innovation in methods of mapping genetic information across layers of cellular state. If I can predict one modality from another, I may expand my understanding of the rules governing these complex regulatory processes.\n\nI will focus on three tasks i.e.,\n\n**Task 1: Modality Prediction**\n\nPredicting the flow of information from DNA to RNA and RNA to Protein Experimental techniques to measure multiple modalities within the same single cell are increasingly becoming available. The demand for these measurements is driven by the promise to provide a deeper insight into the state of a cell.\n\n**Task 2: Modality Matching**\n\nMatching profiles of each cell from different modalities. While joint profiling of two modalities in the same single cell is now possible, most single-cell datasets that exist measure only a single modality.\n\n**Task 3: Joint Embedding**\n\nLearn a joint embedding from multiple modalities The functioning of organs, tissues, and whole organisms is determined by the interplay of cells. Cells are characterised into broad types, which in turn can take on different states.\n\n**Why are you trying it?**\n\nMy task is to predict the labels corresponding to the inputs in the test set. To facilitate submission scoring, I only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\n\nIn the past decade, the advent of single-cell genomics has enabled the measurement of DNA, RNA, and proteins in single cells. These technologies allow the study of biology at an unprecedented scale and resolution. Among the outcomes have been detailed maps of early human embryonic development, the discovery of new disease-associated cell types, and cell-targeted therapeutic interventions. Moreover, with recent advances in experimental techniques it is now possible to measure multiple genomic modalities in the same cell.\n\nWhile multimodal single-cell data is increasingly available, data analysis methods are still scarce. Due to the small volume of a single cell, measurements are sparse and noisy. Differences in molecular sampling depths between cells (sequencing depth) and technical effects from handling cells in batches (batch effects) can often overwhelm biological differences. When analyzing multimodal data, one must account for different feature spaces, as well as shared and unique variation between modalities and between batches. Furthermore, current pipelines for single-cell data analysis treat cells as static snapshots, even when there is an underlying dynamical biological process. Accounting for temporal dynamics alongside state changes over time is an open challenge in single-cell data science.\n\nThere are approximately 37 trillion cells in the human body, all with different behaviors and functions. Understanding how a single genome gives rise to a diversity of cellular states is the key to gaining mechanistic insight into how tissues function or malfunction in health and disease. You can help solve this fundamental challenge for single-cell biology. Being able to solve the prediction problems over time may yield new insights into how gene regulation influences differentiation as blood and immune cells mature.\n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:33.991871Z","iopub.execute_input":"2022-11-03T14:30:33.992593Z","iopub.status.idle":"2022-11-03T14:30:34.062498Z","shell.execute_reply.started":"2022-11-03T14:30:33.992481Z","shell.execute_reply":"2022-11-03T14:30:34.061365Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:34.064288Z","iopub.execute_input":"2022-11-03T14:30:34.064727Z","iopub.status.idle":"2022-11-03T14:30:34.069374Z","shell.execute_reply.started":"2022-11-03T14:30:34.064692Z","shell.execute_reply":"2022-11-03T14:30:34.068240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport copy\nimport gc\nimport math\nimport itertools\nimport pickle\nimport glob\nimport joblib\nimport json\nimport random\nimport re\nimport operator\nimport collections","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:34.070893Z","iopub.execute_input":"2022-11-03T14:30:34.071585Z","iopub.status.idle":"2022-11-03T14:30:34.103499Z","shell.execute_reply.started":"2022-11-03T14:30:34.071532Z","shell.execute_reply":"2022-11-03T14:30:34.102667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch\nimport torch.nn as nn","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:34.106032Z","iopub.execute_input":"2022-11-03T14:30:34.106612Z","iopub.status.idle":"2022-11-03T14:30:35.854187Z","shell.execute_reply.started":"2022-11-03T14:30:34.106579Z","shell.execute_reply":"2022-11-03T14:30:35.853263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import plotly.express as px","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:35.855750Z","iopub.execute_input":"2022-11-03T14:30:35.856258Z","iopub.status.idle":"2022-11-03T14:30:37.824928Z","shell.execute_reply.started":"2022-11-03T14:30:35.856222Z","shell.execute_reply":"2022-11-03T14:30:37.823943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import scipy","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:37.826630Z","iopub.execute_input":"2022-11-03T14:30:37.826999Z","iopub.status.idle":"2022-11-03T14:30:37.833672Z","shell.execute_reply.started":"2022-11-03T14:30:37.826959Z","shell.execute_reply":"2022-11-03T14:30:37.831431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sklearn\nimport sklearn.cluster\nimport sklearn.preprocessing\nimport copy","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:37.835054Z","iopub.execute_input":"2022-11-03T14:30:37.835970Z","iopub.status.idle":"2022-11-03T14:30:38.280064Z","shell.execute_reply.started":"2022-11-03T14:30:37.835933Z","shell.execute_reply":"2022-11-03T14:30:38.278746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from collections import defaultdict\nfrom operator import itemgetter, attrgetter\nfrom tqdm.notebook import tqdm\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import KFold","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.281753Z","iopub.execute_input":"2022-11-03T14:30:38.282095Z","iopub.status.idle":"2022-11-03T14:30:38.288312Z","shell.execute_reply.started":"2022-11-03T14:30:38.282058Z","shell.execute_reply":"2022-11-03T14:30:38.286861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def partial_correlation_score_torch_faster(y_true, y_pred):\n    y_true_centered = y_true - torch.mean(y_true, dim=1)[:,None]\n    y_pred_centered = y_pred - torch.mean(y_pred, dim=1)[:,None]\n    cov_tp = torch.sum(y_true_centered*y_pred_centered, dim=1)/(y_true.shape[1]-1)\n    var_t = torch.sum(y_true_centered**2, dim=1)/(y_true.shape[1]-1)\n    var_p = torch.sum(y_pred_centered**2, dim=1)/(y_true.shape[1]-1)\n    return cov_tp/torch.sqrt(var_t*var_p)\n\ndef correl_loss(pred, tgt):\n    return -torch.mean(partial_correlation_score_torch_faster(tgt, pred))","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.290773Z","iopub.execute_input":"2022-11-03T14:30:38.291242Z","iopub.status.idle":"2022-11-03T14:30:38.302081Z","shell.execute_reply.started":"2022-11-03T14:30:38.291199Z","shell.execute_reply":"2022-11-03T14:30:38.300472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TorchCSR = collections.namedtuple(\"TrochCSR\", \"data indices indptr shape\")\n\ndef load_csr_data_to_gpu(train_inputs):\n    th_data = torch.from_numpy(train_inputs.data).to(device)\n    th_indices = torch.from_numpy(train_inputs.indices).to(device)\n    th_indptr = torch.from_numpy(train_inputs.indptr).to(device)\n    th_shape = train_inputs.shape\n    return TorchCSR(th_data, th_indices, th_indptr, th_shape)\n\ndef make_coo_batch(torch_csr, indx):\n    th_data, th_indices, th_indptr, th_shape = torch_csr\n    start_pts = th_indptr[indx]\n    end_pts = th_indptr[indx+1]\n    coo_data = torch.cat([th_data[start_pts[i]: end_pts[i]] for i in range(len(start_pts))], dim=0)\n    coo_col = torch.cat([th_indices[start_pts[i]: end_pts[i]] for i in range(len(start_pts))], dim=0)\n    coo_row = torch.repeat_interleave(torch.arange(indx.shape[0], device=device), th_indptr[indx+1] - th_indptr[indx])\n    coo_batch = torch.sparse_coo_tensor(torch.vstack([coo_row, coo_col]), coo_data, [indx.shape[0], th_shape[1]])\n    return coo_batch\n\n\ndef make_coo_batch_slice(torch_csr, start, end):\n    th_data, th_indices, th_indptr, th_shape = torch_csr\n    if end > th_shape[0]:\n        end = th_shape[0]\n    start_pts = th_indptr[start]\n    end_pts = th_indptr[end]\n    coo_data = th_data[start_pts: end_pts]\n    coo_col = th_indices[start_pts: end_pts]\n    coo_row = torch.repeat_interleave(torch.arange(end-start, device=device), th_indptr[start+1:end+1] - th_indptr[start:end])\n    coo_batch = torch.sparse_coo_tensor(torch.vstack([coo_row, coo_col]), coo_data, [end-start, th_shape[1]])\n    return coo_batch","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.308273Z","iopub.execute_input":"2022-11-03T14:30:38.308600Z","iopub.status.idle":"2022-11-03T14:30:38.324112Z","shell.execute_reply.started":"2022-11-03T14:30:38.308563Z","shell.execute_reply":"2022-11-03T14:30:38.322899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class DataLoaderCOO:\n    def __init__(self, train_inputs, train_targets, train_idx=None, \n                 *,\n                batch_size=512, shuffle=False, drop_last=False):\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.drop_last = drop_last\n        \n        self.train_inputs = train_inputs\n        self.train_targets = train_targets\n        \n        self.train_idx = train_idx\n        \n        self.nb_examples = len(self.train_idx) if self.train_idx is not None else train_inputs.shape[0]\n        \n        self.nb_batches = self.nb_examples//batch_size\n        if not drop_last and not self.nb_examples%batch_size==0:\n            self.nb_batches +=1\n        \n    def __iter__(self):\n        if self.shuffle:\n            shuffled_idx = torch.randperm(self.nb_examples, device=device)\n            if self.train_idx is not None:\n                idx_array = self.train_idx[shuffled_idx]\n            else:\n                idx_array = shuffled_idx\n        else:\n            if self.train_idx is not None:\n                idx_array = self.train_idx\n            else:\n                idx_array = None\n            \n        for i in range(self.nb_batches):\n            slc = slice(i*self.batch_size, (i+1)*self.batch_size)\n            if idx_array is None:\n                inp_batch = make_coo_batch_slice(self.train_inputs, i*self.batch_size, (i+1)*self.batch_size)\n                if self.train_targets is None:\n                    tgt_batch = None\n                else:\n                    tgt_batch = make_coo_batch_slice(self.train_targets, i*self.batch_size, (i+1)*self.batch_size)\n            else:\n                idx_batch = idx_array[slc]\n                inp_batch = make_coo_batch(self.train_inputs, idx_batch)\n                if self.train_targets is None:\n                    tgt_batch = None\n                else:\n                    tgt_batch = make_coo_batch(self.train_targets, idx_batch)\n            yield inp_batch, tgt_batch\n            \n            \n    def __len__(self):\n        return self.nb_batches","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.327843Z","iopub.execute_input":"2022-11-03T14:30:38.328215Z","iopub.status.idle":"2022-11-03T14:30:38.344941Z","shell.execute_reply.started":"2022-11-03T14:30:38.328187Z","shell.execute_reply":"2022-11-03T14:30:38.343812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class MLP(nn.Module):\n    def __init__(self, layer_size_lst, add_final_activation=False):\n        super().__init__()\n        \n        assert len(layer_size_lst) > 2\n        \n        layer_lst = []\n        for i in range(len(layer_size_lst)-1):\n            sz1 = layer_size_lst[i]\n            sz2 = layer_size_lst[i+1]\n            layer_lst += [nn.Linear(sz1, sz2)]\n            if i != len(layer_size_lst)-2 or add_final_activation:\n                 layer_lst += [nn.ReLU()]\n        self.mlp = nn.Sequential(*layer_lst)\n        \n    def forward(self, x):\n        return self.mlp(x)\n    \ndef build_model():\n    model = MLP([INPUT_SIZE] + config[\"layers\"] + [OUTPUT_SIZE])\n    if config[\"head\"] == \"softplus\":\n        model = nn.Sequential(model, nn.Softplus())\n    else:\n        assert config[\"head\"] is None\n    return model","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.348995Z","iopub.execute_input":"2022-11-03T14:30:38.349338Z","iopub.status.idle":"2022-11-03T14:30:38.360398Z","shell.execute_reply.started":"2022-11-03T14:30:38.349310Z","shell.execute_reply":"2022-11-03T14:30:38.359181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def test_fn_ensemble(model_list, dl_test):\n\n    res = torch.empty(\n        (dl_test.nb_examples, OUTPUT_SIZE), \n        device=device, dtype=torch.float32)\n    \n    for model in model_list:\n        model.eval()\n        \n    cur = 0\n    for inpt, tgt in tqdm(dl_test):\n        mb_size = inpt.shape[0]\n\n        with torch.no_grad():\n            pred_list = []\n            for model in model_list:\n                pred = model(inpt)\n                pred_list.append(pred)\n            pred = sum(pred_list)/len(pred_list)\n            \n        res[cur:cur+pred.shape[0]] = pred\n        cur += pred.shape[0]\n            \n    return {\"preds\":res}","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.363810Z","iopub.execute_input":"2022-11-03T14:30:38.364081Z","iopub.status.idle":"2022-11-03T14:30:38.374319Z","shell.execute_reply.started":"2022-11-03T14:30:38.364057Z","shell.execute_reply":"2022-11-03T14:30:38.373325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if torch.cuda.is_available():\n    device = torch.device(\"cuda:0\")\n    print(f\"machine has {torch.cuda.device_count()} cuda devices\")\n    print(f\"model of first cuda device is {torch.cuda.get_device_name(0)}\")\nelse:\n    device = torch.device(\"cpu\")","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.375877Z","iopub.execute_input":"2022-11-03T14:30:38.376208Z","iopub.status.idle":"2022-11-03T14:30:38.448795Z","shell.execute_reply.started":"2022-11-03T14:30:38.376173Z","shell.execute_reply":"2022-11-03T14:30:38.447536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"INPUT_SIZE = 228942 \nOUTPUT_SIZE = 23418","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.452301Z","iopub.execute_input":"2022-11-03T14:30:38.453093Z","iopub.status.idle":"2022-11-03T14:30:38.457721Z","shell.execute_reply.started":"2022-11-03T14:30:38.453058Z","shell.execute_reply":"2022-11-03T14:30:38.456607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_inputs = np.load(\"../input/msci-multiome-torch-quickstart-w-sparse-tensors/max_inputs.npz\")[\"max_inputs\"]\nmax_inputs = torch.from_numpy(max_inputs)[0].to(device)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:38.459397Z","iopub.execute_input":"2022-11-03T14:30:38.459973Z","iopub.status.idle":"2022-11-03T14:30:41.657182Z","shell.execute_reply.started":"2022-11-03T14:30:38.459938Z","shell.execute_reply":"2022-11-03T14:30:41.656098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntest_inputs = scipy.sparse.load_npz(\n    \"../input/multimodal-single-cell-as-sparse-matrix/test_multi_inputs_values.sparse.npz\")","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:30:41.659955Z","iopub.execute_input":"2022-11-03T14:30:41.660752Z","iopub.status.idle":"2022-11-03T14:31:13.328768Z","shell.execute_reply.started":"2022-11-03T14:30:41.660711Z","shell.execute_reply":"2022-11-03T14:31:13.327627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntest_inputs = load_csr_data_to_gpu(test_inputs)\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:13.330074Z","iopub.execute_input":"2022-11-03T14:31:13.330891Z","iopub.status.idle":"2022-11-03T14:31:14.082140Z","shell.execute_reply.started":"2022-11-03T14:31:13.330853Z","shell.execute_reply":"2022-11-03T14:31:14.080985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_inputs.data[...] /= max_inputs[test_inputs.indices.long()]","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:14.083901Z","iopub.execute_input":"2022-11-03T14:31:14.084286Z","iopub.status.idle":"2022-11-03T14:31:14.110102Z","shell.execute_reply.started":"2022-11-03T14:31:14.084249Z","shell.execute_reply":"2022-11-03T14:31:14.109234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.max(test_inputs.data)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:14.111598Z","iopub.execute_input":"2022-11-03T14:31:14.111956Z","iopub.status.idle":"2022-11-03T14:31:14.139929Z","shell.execute_reply.started":"2022-11-03T14:31:14.111922Z","shell.execute_reply":"2022-11-03T14:31:14.138839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_list = []\nfor fn in tqdm(glob.glob(\"../input/msci-multiome-torch-quickstart-w-sparse-tensors/*_best_params.pth\")):\n    prefix = fn[:-len(\"_best_params.pth\")]\n    config_fn = prefix + \"_config.pkl\"\n    \n    config = pickle.load(open(config_fn, \"rb\"))\n    \n    model = build_model() \n    model.to(device)\n    \n    params = torch.load(fn)\n    model.load_state_dict(params)\n    \n    model_list.append(model)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:14.141708Z","iopub.execute_input":"2022-11-03T14:31:14.142098Z","iopub.status.idle":"2022-11-03T14:31:22.200234Z","shell.execute_reply.started":"2022-11-03T14:31:14.142059Z","shell.execute_reply":"2022-11-03T14:31:22.199350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dl_test = DataLoaderCOO(test_inputs, None, train_idx=None,\n                batch_size=512, shuffle=False, drop_last=False)","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:22.203338Z","iopub.execute_input":"2022-11-03T14:31:22.203637Z","iopub.status.idle":"2022-11-03T14:31:22.209214Z","shell.execute_reply.started":"2022-11-03T14:31:22.203610Z","shell.execute_reply":"2022-11-03T14:31:22.207422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_pred = test_fn_ensemble(model_list, dl_test)[\"preds\"]","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:22.210734Z","iopub.execute_input":"2022-11-03T14:31:22.211286Z","iopub.status.idle":"2022-11-03T14:31:45.192879Z","shell.execute_reply.started":"2022-11-03T14:31:22.211246Z","shell.execute_reply":"2022-11-03T14:31:45.191905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del model_list\ndel dl_test\ndel test_inputs\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:45.194574Z","iopub.execute_input":"2022-11-03T14:31:45.195298Z","iopub.status.idle":"2022-11-03T14:31:45.333541Z","shell.execute_reply.started":"2022-11-03T14:31:45.195260Z","shell.execute_reply":"2022-11-03T14:31:45.332621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_pred.shape","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:45.335385Z","iopub.execute_input":"2022-11-03T14:31:45.335860Z","iopub.status.idle":"2022-11-03T14:31:45.348044Z","shell.execute_reply.started":"2022-11-03T14:31:45.335807Z","shell.execute_reply":"2022-11-03T14:31:45.346962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\neval_ids = pd.read_parquet(\"../input/multimodal-single-cell-as-sparse-matrix/evaluation.parquet\")\neval_ids.cell_id = eval_ids.cell_id.astype(pd.CategoricalDtype())\neval_ids.gene_id = eval_ids.gene_id.astype(pd.CategoricalDtype())","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:31:45.349829Z","iopub.execute_input":"2022-11-03T14:31:45.350196Z","iopub.status.idle":"2022-11-03T14:32:18.104380Z","shell.execute_reply.started":"2022-11-03T14:31:45.350160Z","shell.execute_reply":"2022-11-03T14:32:18.103244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.Series(name='target',\n                       index=pd.MultiIndex.from_frame(eval_ids), \n                       dtype=np.float32)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-11-03T14:32:18.106603Z","iopub.execute_input":"2022-11-03T14:32:18.107294Z","iopub.status.idle":"2022-11-03T14:32:37.624283Z","shell.execute_reply.started":"2022-11-03T14:32:18.107252Z","shell.execute_reply":"2022-11-03T14:32:37.623347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Did it work?**\n\nMy task is to predict the labels corresponding to the inputs in the test set. To facilitate submission scoring, I only require predictions on a subset of the Multiome data. This subset was created by sampling 30% of the Multiome rows, and for each row, 15% of the columns. The sample of columns varies from row-to-row. All of the CITEseq labels are scored.\n\nThe data splits are arranged as follows:\n\nThe training set comprises samples only from donors 13176, 31800, and 32606. The public test set comprises samples only from donor 27678. The private test set comprises samples from all four donors.\n\nFor the Multiome samples, the training set comprises samples only from days 2, 3, 4, and 7. The public test set comprises samples only from days 2, 3, and 7. The private test set comprises data only from day 10.\n\nFor the CITEseq samples, the training set comprises samples only from days 2, 3, and 4. The public test set also comprises samples only from days 2, 3, and 4. The private test set comprises samples only from day 7. There are no day 10 CITEseq samples in any split.\n\n**What did you not understand about it?**\n\nIn the test set, taken from an unseen later time point in the dataset, competitors will be provided with one modality and be tasked with predicting a paired modality measured in the same cell. The added challenge of this competition is that the test data will be from a later time point than any time point in the training data.\n\nThe dataset for this competition comprises single-cell multiomics data collected from mobilized peripheral CD34+ hematopoietic stem and progenitor cells (HSPCs) isolated from four healthy human donors. More information about the cells can be found on the vendor website.\n\n**Share your ideas to accomplish the goals !!!**","metadata":{}}]}