{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":51294,"databundleVersionId":6923401,"sourceType":"competition"},{"sourceId":7107895,"sourceType":"datasetVersion","datasetId":4098015}],"dockerImageVersionId":30615,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Description\nThat's an edit from the starter kit by Iafoss (https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/440702)\n\nMy skills with torch and transformers were inexistent so I took his notebook as tutorial. Thanks to Iafoss and other users in that notebook for what I've learned in this competition.\n\nI had no idea of RNA prediction, so I searched on google and found the work of Shujun He, Baizhen Gao, Rushant Sabnis, and Qing Sun (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9851316/) and I started experimenting with pseudoknots like many of us. I've been impressed by the boost in my predictions. In the same work, I've tried to make a custom SelfAttentionHead to incorporate bpp, but I've failed. I must admit that the fact that I improperly used torch.repeat() in order to update the different heads could have been the principal reason.\nDisappointed by my results on this approximation, I've changed to a simplest pov and implemented the bpp as an embedding for the starting vectors. So by applying a convolution to average the bpp with the distance matrix plus a 5th channel for the diagonal, we would obtain a matrix of scores to turn each of the initial 4 vocabulary vectors into an embedding that would average all the words on the sequence depending on their corresponding bpp, distances and themselves in a way that the model would learn by itself. My hypothesis was that such embedding would help with generalization too. At least for public LB I've achieved good results. Better than the ones I've obtained with pseudoknots with only raw initial sequences and bpp.\n\nFinal classification 93/755 with 0.17894, proud of myself :/\n\nThe thing is that by looking discarted submissions one obtained 0.15595/0.15599, very similar private/public. Was an slices aproximation that I've been working on thinking on generalization. But I left behind a month ago. I'll try to recover the approach. But now I need to sleep. ","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport os, gc\nimport numpy as np\nfrom sklearn.model_selection import KFold\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nfrom fastai.vision.all import *","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.execute_input":"2023-11-22T09:08:26.351339Z","iopub.status.busy":"2023-11-22T09:08:26.350717Z","iopub.status.idle":"2023-11-22T09:08:26.358758Z","shell.execute_reply":"2023-11-22T09:08:26.357320Z","shell.execute_reply.started":"2023-11-22T09:08:26.351280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\n    torch.backends.cudnn.benchmark = True","metadata":{"_kg_hide-input":true,"execution":{"iopub.execute_input":"2023-11-22T09:08:35.580078Z","iopub.status.busy":"2023-11-22T09:08:35.579216Z","iopub.status.idle":"2023-11-22T09:08:35.586576Z","shell.execute_reply":"2023-11-22T09:08:35.585385Z","shell.execute_reply.started":"2023-11-22T09:08:35.580032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fname = 'example0'\nPATH = '/kaggle/input/stanford-ribonanza-rna-folding-converted/'\nOUT = './'\nbs = 32\nnum_workers = 0\nSEED = 2023\nnfolds = 4\ndevice = 'cuda' if torch.cuda.is_available() else 'cpu'","metadata":{"execution":{"iopub.execute_input":"2023-11-22T09:08:39.194036Z","iopub.status.busy":"2023-11-22T09:08:39.193607Z","iopub.status.idle":"2023-11-22T09:08:39.201505Z","shell.execute_reply":"2023-11-22T09:08:39.199941Z","shell.execute_reply.started":"2023-11-22T09:08:39.194002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data\n\nI've tried to amplify my train sequences by defining weighted custom losses depending on the reactivity error. So mistakes with high error positions would contribute less to the loss, but without success. At the end, I've chosen to train with a smallest but confident set of sequences. ","metadata":{}},{"cell_type":"code","source":"seed_everything(SEED)\nos.makedirs(OUT, exist_ok=True)\ndf = pd.read_csv(\"/kaggle/input/stanford-ribonanza-rna-folding/train_data.csv\")\nreact = [c for c in df.columns if 'reactivity_0' in c]\nerror = [c for c in df.columns if 'reactivity_e' in c]\ndf_2A3 = df.loc[df.experiment_type=='2A3_MaP']\ndf_DMS = df.loc[df.experiment_type=='DMS_MaP']\ndel df\nm = (df_2A3['SN_filter'].values > 0) & (df_DMS['SN_filter'].values > 0)\ndf_2A3 = df_2A3.loc[m].reset_index(drop=True)\ndf_DMS = df_DMS.loc[m].reset_index(drop=True)\ndf = df_2A3[['sequence_id','sequence','signal_to_noise'] + react + error].join(df_DMS[['signal_to_noise'] + react + error],lsuffix='_2A3',rsuffix='_DMS')\ndel df_2A3, df_DMS\ndf_bpp = pd.read_csv(\"/kaggle/input/bpp-files-corrected/bpp.csv\")\nprint('len before merge: ',len(df))\ndf = pd.merge(df, df_bpp, on=\"sequence_id\")\nprint('len after merge: ',len(df))\nbpp_root_dir = \"/kaggle/input/stanford-ribonanza-rna-folding/Ribonanza_bpp_files\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class RNA_Dataset(Dataset):\n    def __init__(self, df, mode='train', seed=2023, fold=0, nfolds=4, \n                 mask_only=False, **kwargs):\n        self.seq_map = {'A':0,'C':1,'G':2,'U':3}\n        self.Lmax = 206\n        df['L'] = df.sequence.apply(len)\n        \n        split = list(KFold(n_splits=nfolds, random_state=seed, \n                shuffle=True).split(df))[fold][0 if mode=='train' else 1]\n        df = df.iloc[split].reset_index(drop=True)\n        \n        self.seq = df['sequence'].values\n        self.L = df['L'].values\n        self.filepaths = df['path'].values\n        self.sequence_id = df['sequence_id'].values\n        self.bpp_root_dir = bpp_root_dir\n        \n        self.react_2A3 = df[[c + '_2A3' for c in react]].values\n        self.react_DMS = df[[c + '_DMS' for c in react]].values\n#       self.react_err_2A3 = df[[c + '_2A3' for c in error]].values\n#       self.react_err_DMS = df[[c + '_DMS' for c in error]].values\n#       self.sn_2A3 = df['signal_to_noise_2A3'].values\n#       self.sn_DMS = df['signal_to_noise_DMS'].values\n        self.mask_only = mask_only\n        \n    def __len__(self):\n        return len(self.seq)  \n    \n    def __getitem__(self, idx):\n        L = self.L[idx]\n        Lmax = self.Lmax\n\n        seq = self.seq[idx]\n        if self.mask_only:\n            mask = torch.zeros(self.Lmax, dtype=torch.bool)\n            mask[:L] = True\n            return {'mask':mask},{'mask':mask}\n        seq = [self.seq_map[s] for s in seq]\n        seq = np.array(seq)\n        mask = torch.zeros(self.Lmax, dtype=torch.bool)\n        mask[:L] = True\n        seq = np.pad(seq,(0,self.Lmax-L))\n\n        filepath = self.bpp_root_dir + '/' + self.filepaths[idx] + '/' + self.sequence_id[idx] + '.txt'\n        df_bpp = pd.read_csv(filepath,header=None,sep=' ')\n        df_bpp = pd.concat((df_bpp,pd.DataFrame({0:df_bpp[1],1:df_bpp[0],2:df_bpp[2]})))# Adding transpose indices.\n        df_bpp.loc[:,[0,1]] -= 1# Correcting indices.\n        indices = (df_bpp[[0,1]].values).swapaxes(0,1)\n        values = df_bpp[2].values\n        bpp = torch.sparse_coo_tensor(indices, values, [Lmax, Lmax],dtype=torch.float32).to_dense()\n\n        react = torch.from_numpy(np.stack([self.react_2A3[idx],\n                                           self.react_DMS[idx]],-1))\n#       react_err = torch.from_numpy(np.stack([self.react_err_2A3[idx],\n#                                              self.react_err_DMS[idx]],-1))\n#       sn = torch.FloatTensor([self.sn_2A3[idx],self.sn_DMS[idx]])\n#       \n        return {'seq':torch.from_numpy(seq), 'bpp':bpp, 'mask':mask}, \\\n               {'react':react.float(), 'mask':mask}\n    \nclass LenMatchBatchSampler(torch.utils.data.BatchSampler):\n    def __iter__(self):\n        buckets = [[]] * 100\n        yielded = 0\n\n        for idx in self.sampler:\n            s = self.sampler.data_source[idx]\n            if isinstance(s,tuple): L = s[0][\"mask\"].sum()\n            else: L = s[\"mask\"].sum()\n            L = max(1,L // 16) \n            if len(buckets[L]) == 0:  buckets[L] = []\n            buckets[L].append(idx)\n            \n            if len(buckets[L]) == self.batch_size:\n                batch = list(buckets[L])\n                yield batch\n                yielded += 1\n                buckets[L] = []\n                \n        batch = []\n        leftover = [idx for bucket in buckets for idx in bucket]\n\n        for idx in leftover:\n            batch.append(idx)\n            if len(batch) == self.batch_size:\n                yielded += 1\n                yield batch\n                batch = []\n\n        if len(batch) > 0 and not self.drop_last:\n            yielded += 1\n            yield batch\n            \ndef dict_to(x, device='cuda'):\n    return {k:x[k].to(device) for k in x}\n\ndef to_device(x, device='cuda'):\n    return tuple(dict_to(e,device) for e in x)\n\nclass DeviceDataLoader:\n    def __init__(self, dataloader, device='cuda'):\n        self.dataloader = dataloader\n        self.device = device\n    \n    def __len__(self):\n        return len(self.dataloader)\n    \n    def __iter__(self):\n        for batch in self.dataloader:\n            yield tuple(dict_to(x, self.device) for x in batch)","metadata":{"execution":{"iopub.execute_input":"2023-10-20T16:49:37.055698Z","iopub.status.busy":"2023-10-20T16:49:37.054814Z","iopub.status.idle":"2023-10-20T16:49:37.083653Z","shell.execute_reply":"2023-10-20T16:49:37.082811Z","shell.execute_reply.started":"2023-10-20T16:49:37.055658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model\nDue to my modest resources and after some experiments, I've found that the best I could afford was the original basic transformer, but with double layers.","metadata":{}},{"cell_type":"code","source":"class SinusoidalPosEmb(nn.Module):\n    def __init__(self, dim=16, M=10000):\n        super().__init__()\n        self.dim = dim\n        self.M = M\n\n    def forward(self, x):\n        device = x.device\n        half_dim = self.dim // 2\n        emb = math.log(self.M) / half_dim\n        emb = torch.exp(torch.arange(half_dim, device=device) * (-emb))\n        emb = x[...,None] * emb[None,...]\n        emb = torch.cat((emb.sin(), emb.cos()), dim=-1)\n        return emb\n\nclass RNA_Model(nn.Module):\n    def __init__(self, dim=192, depth=12, head_size=32, **kwargs):\n        super().__init__()\n        self.emb = nn.Embedding(4,dim)\n#       Convolution to the bpp + Distancesx3 + Diagonal channels of scores for the embedding transformation\n#       A single convolution, I want this step simple for a fast convergence\n        self.conv = nn.Conv2d(5,1,1)\n        self.pos_enc = SinusoidalPosEmb(dim)\n        self.transformer = nn.TransformerEncoder(\n                nn.TransformerEncoderLayer(d_model=dim, nhead=dim//head_size, dim_feedforward=4*dim,\n                dropout=0.1, activation=nn.GELU(), batch_first=True, norm_first=True, device=device), depth)\n        self.proj_out = nn.Linear(dim,2)\n    \n    def forward(self, x0):\n        mask = x0['mask']\n        L = mask.sum(-1).max()\n        mask = mask[:,:L]\n        x = x0['seq'][:,:L]\n#       Reinforced scores as described in:\n#       RNAdegformer: accurate prediction of mRNA degradation at nucleotide resolution with deep learning\n#       Shujun He, Baizhen Gao, Rushant Sabnis and Qing Sun\n#       Corresponding author. Qing Sun, Department of Chemical Engineering, Texas A&M University, 100 Spence St., 77843 TX, USA. Tel.: 979-845-3401;\n#       E-mail: sunqing@tamu.edu\n#       BPP\n        bpp = x0['bpp'][:,:L,:L]\n        scr = torch.zeros(list(bpp.shape)+[5],dtype=torch.float32,device=device)\n        scr[:,:,:,0] = bpp\n#       Distance matrix\n#         0   1 1/2 1/3 ...\n#         1   0   1 1/2 ...\n#       1/2   1   0   1 ...\n#       1/3 1/2   1   0 ...\n#       ... ... ... ... ...\n        distances = torch.arange(1,L)\n        distances[1:] = 1/distances[1:]\n        distance_mat = torch.zeros((L,L,3),dtype=torch.float32)\n        for i in range(L):\n            dist = distances[:L-i-1]\n            distance_mat[i,i+1:L,0] = dist\n            distance_mat[i+1:L,i,0] = dist\n\n        distance_mat[:L,:L,1] = distance_mat[:L,:L,0]*distance_mat[:L,:L,0]# Squares\n        distance_mat[:L,:L,2] = distance_mat[:L,:L,1]*distance_mat[:L,:L,0]# Cubes\n\n        scr[:,:,:,1:4] = distance_mat\n#       Diagonal\n        scr[:,:,:,-1] = torch.diag(torch.ones(L,dtype=torch.float32))\n\n        scr = torch.swapaxes(scr,1,-1)\n        bpp = self.conv(scr).squeeze(1)\n#       bpp = nn.ReLU()(bpp)\n#       bpp = nn.Softmax()(bpp)\n        \n        pos = torch.arange(L, device=x.device).unsqueeze(0)\n        pos = self.pos_enc(pos)\n        x = self.emb(x)\n        x = torch.matmul(bpp,x)\n        x = x + pos\n        \n        x = self.transformer(x,src_key_padding_mask=~mask)\n        x = self.proj_out(x)\n        \n        return x","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loss & Metric\nNo changes here.","metadata":{}},{"cell_type":"code","source":"def loss(pred,target):\n    p = pred[target['mask'][:,:pred.shape[1]]]\n    y = target['react'][target['mask']].clip(0,1)\n    loss = F.l1_loss(p, y, reduction='none')\n    loss = loss[~torch.isnan(loss)].mean()\n    \n    return loss\n\nclass MAE(Metric):\n    def __init__(self): \n        self.reset()\n        \n    def reset(self): \n        self.x,self.y = [],[]\n        \n    def accumulate(self, learn):\n        x = learn.pred[learn.y['mask'][:,:learn.pred.shape[1]]]\n        y = learn.y['react'][learn.y['mask']].clip(0,1)\n        self.x.append(x)\n        self.y.append(y)\n\n    @property\n    def value(self):\n        x,y = torch.cat(self.x,0),torch.cat(self.y,0)\n        loss = F.l1_loss(x, y, reduction='none')\n        loss = loss[~torch.isnan(loss)].mean()\n        return loss","metadata":{"execution":{"iopub.execute_input":"2023-10-20T16:49:55.236238Z","iopub.status.busy":"2023-10-20T16:49:55.235289Z","iopub.status.idle":"2023-10-20T16:49:55.246312Z","shell.execute_reply":"2023-10-20T16:49:55.245244Z","shell.execute_reply.started":"2023-10-20T16:49:55.236196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training\nI've achieved the best results with trainings on batches of 32 with 1e-3 of learning rate.","metadata":{}},{"cell_type":"code","source":"bs = 32\nfor fold in []: # running multiple folds at kaggle may cause OOM\n    ds_train = RNA_Dataset(df, mode='train', fold=fold, nfolds=nfolds)\n    ds_train_len = RNA_Dataset(df, mode='train', fold=fold, \n                nfolds=nfolds, mask_only=True)\n    sampler_train = torch.utils.data.RandomSampler(ds_train_len)\n    len_sampler_train = LenMatchBatchSampler(sampler_train, batch_size=bs,\n                drop_last=True)\n    dl_train = DeviceDataLoader(torch.utils.data.DataLoader(ds_train, \n                batch_sampler=len_sampler_train, num_workers=num_workers,\n                persistent_workers=False), device)\n\n    ds_val = RNA_Dataset(df, mode='eval', fold=fold, nfolds=nfolds)\n    ds_val_len = RNA_Dataset(df, mode='eval', fold=fold, nfolds=nfolds, \n               mask_only=True)\n    sampler_val = torch.utils.data.SequentialSampler(ds_val_len)\n    len_sampler_val = LenMatchBatchSampler(sampler_val, batch_size=bs, \n               drop_last=False)\n    dl_val= DeviceDataLoader(torch.utils.data.DataLoader(ds_val, \n               batch_sampler=len_sampler_val, num_workers=num_workers), device)\n    gc.collect()\n\n    data = DataLoaders(dl_train,dl_val)\n    model = RNA_Model(dim=192, depth=24, head_size=32)   \n    model = model.to(device)\n    learn = Learner(data, model, loss_func=loss,cbs=[GradientClip(3.0),SaveModelCallback (monitor='train_loss', comp=None, min_delta=0.0,\n                    fname='model', every_epoch=True, at_end=False,\n                    with_opt=True, reset_on_fit=True)],\n                metrics=[MAE()])#.to_fp16() \n    #fp16 doesn't help at P100 but gives x1.6-1.8 speedup at modern hardware\n\n    learn.fit_one_cycle(32, lr_max=1e-3, wd=0.05, pct_start=0.02)\n    torch.save(learn.model.state_dict(),os.path.join(OUT,f'{fname}_{fold}.pth'))\n    gc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Results from the different folds.","metadata":{}},{"cell_type":"code","source":"bpp_24_3 = '''0\t0.176180\t0.178217\t0.178550\t50:03\n1\t0.158920\t0.157980\t0.158181\t47:38\n2\t0.150089\t0.149436\t0.149730\t47:50\n3\t0.147765\t0.146751\t0.147066\t43:52\n4\t0.144205\t0.143995\t0.144368\t43:40\n5\t0.143693\t0.141611\t0.142035\t43:03\n6\t0.141300\t0.141435\t0.141864\t46:09\n7\t0.139334\t0.138799\t0.139249\t48:02\n8\t0.139512\t0.137960\t0.138440\t45:28\n9\t0.137833\t0.137699\t0.138210\t45:11\n10\t0.136555\t0.136052\t0.136579\t45:20\n11\t0.136631\t0.135909\t0.136471\t45:30\n12\t0.135671\t0.134839\t0.135403\t47:41\n13\t0.133340\t0.133684\t0.134257\t48:38\n14\t0.132652\t0.133382\t0.133959\t45:28\n15\t0.133027\t0.132686\t0.133285\t45:34\n16\t0.132554\t0.132143\t0.132744\t46:16\n17\t0.129787\t0.131363\t0.131989\t48:41\n18\t0.129604\t0.130632\t0.131269\t50:26\n19\t0.128786\t0.130233\t0.130876\t49:22\n20\t0.128188\t0.129925\t0.130584\t48:21\n21\t0.126282\t0.128819\t0.129495\t47:59\n22\t0.124807\t0.128654\t0.129339\t48:18\n24\t0.124416\t0.127760\t0.128467\t49:14\n25\t0.122915\t0.127515\t0.128231\t49:40\n26\t0.123242\t0.127431\t0.128152\t47:20\n27\t0.120564\t0.127240\t0.127969\t49:28\n28\t0.120724\t0.126934\t0.127667\t47:17\n29\t0.119394\t0.126950\t0.127687\t43:59\n30\t0.118568\t0.126957\t0.127695\t43:31\n31\t0.120777\t0.126955\t0.127694\t47:36'''","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:38:30.892614Z","iopub.execute_input":"2023-12-08T10:38:30.892895Z","iopub.status.idle":"2023-12-08T10:38:30.918817Z","shell.execute_reply.started":"2023-12-08T10:38:30.892873Z","shell.execute_reply":"2023-12-08T10:38:30.918115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bpp_24_2 = '''0\t0.176537\t0.174162\t0.174479\t50:22\n1\t0.159618\t0.156158\t0.156452\t38:13\n2\t0.152168\t0.149000\t0.149335\t33:46\n3\t0.148887\t0.146146\t0.146529\t34:57\n4\t0.146359\t0.144162\t0.144572\t34:30\n5\t0.143939\t0.142137\t0.142568\t42:30\n6\t0.141378\t0.141006\t0.141473\t46:18\n7\t0.141651\t0.139295\t0.139759\t39:46\n8\t0.141320\t0.137615\t0.138109\t36:22\n9\t0.139408\t0.136922\t0.137443\t36:55\n10\t0.136797\t0.135979\t0.136508\t38:17\n11\t0.136318\t0.134932\t0.135476\t41:27\n12\t0.136849\t0.134358\t0.134915\t50:52\n13\t0.133786\t0.133512\t0.134086\t46:37\n14\t0.133424\t0.132741\t0.133337\t38:30\n15\t0.133346\t0.132787\t0.133386\t40:04\n16\t0.131859\t0.131366\t0.131991\t45:37\n17\t0.130861\t0.131001\t0.131649\t44:34\n18\t0.130448\t0.130450\t0.131084\t46:56\n19\t0.128067\t0.130009\t0.130662\t44:41\n20\t0.128294\t0.128978\t0.129651\t45:21\n21\t0.127871\t0.128649\t0.129337\t46:55\n22\t0.124597\t0.127964\t0.128656\t45:16\n23\t0.123940\t0.127727\t0.128430\t47:21\n24\t0.123450\t0.127405\t0.128117\t48:00\n25\t0.123456\t0.127510\t0.128238\t45:46\n26\t0.122082\t0.127030\t0.127763\t45:56\n27\t0.121361\t0.126878\t0.127614\t46:11\n28\t0.120326\t0.126678\t0.127421\t48:53\n29\t0.119614\t0.126672\t0.127416\t50:07\n30\t0.121019\t0.126573\t0.127319\t50:02\n31\t0.119301\t0.126600\t0.127345\t46:06'''","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:38:43.558976Z","iopub.execute_input":"2023-12-08T10:38:43.559276Z","iopub.status.idle":"2023-12-08T10:38:43.563596Z","shell.execute_reply.started":"2023-12-08T10:38:43.559252Z","shell.execute_reply":"2023-12-08T10:38:43.562728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bpp_24_1 = '''0\t0.174314\t0.174929\t0.175245\t49:59\n1\t0.157779\t0.156823\t0.157146\t34:47\n2\t0.150059\t0.149707\t0.150009\t32:43\n3\t0.147674\t0.145331\t0.145690\t32:37\n4\t0.145308\t0.142871\t0.143311\t32:40\n5\t0.142889\t0.142706\t0.143127\t32:42\n6\t0.141634\t0.141501\t0.141991\t41:38\n7\t0.140680\t0.140180\t0.140657\t45:39\n8\t0.138863\t0.138092\t0.138583\t36:34\n9\t0.137016\t0.138689\t0.139213\t37:43\n10\t0.135704\t0.136201\t0.136764\t34:35\n11\t0.136345\t0.135417\t0.135992\t37:36\n12\t0.135668\t0.134852\t0.135434\t35:03\n13\t0.133904\t0.134210\t0.134795\t38:16\n14\t0.132174\t0.133192\t0.133797\t44:48\n15\t0.131680\t0.132652\t0.133271\t47:00\n16\t0.131378\t0.131804\t0.132434\t36:23\n17\t0.130098\t0.131345\t0.131994\t38:42\n18\t0.128925\t0.130592\t0.131260\t49:15\n19\t0.128459\t0.130030\t0.130712\t45:57\n20\t0.126981\t0.129745\t0.130433\t47:21\n21\t0.126035\t0.129070\t0.129763\t46:36\n22\t0.125090\t0.128903\t0.129617\t41:24\n23\t0.124419\t0.128113\t0.128834\t41:29\n24\t0.124148\t0.127836\t0.128571\t41:04\n25\t0.123795\t0.127598\t0.128339\t43:49\n26\t0.119715\t0.127231\t0.127985\t46:31\n27\t0.120699\t0.127229\t0.127986\t43:58\n28\t0.119039\t0.127167\t0.127929\t46:47\n29\t0.120494\t0.126996\t0.127762\t48:12\n30\t0.119801\t0.126964\t0.127729\t45:46\n31\t0.118476\t0.126972\t0.127738\t48:03'''","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:38:46.045231Z","iopub.execute_input":"2023-12-08T10:38:46.045549Z","iopub.status.idle":"2023-12-08T10:38:46.050119Z","shell.execute_reply.started":"2023-12-08T10:38:46.045523Z","shell.execute_reply":"2023-12-08T10:38:46.049231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bpp_24_0 = '''0\t0.175593\t0.177211\t0.177592\t50:28\n1\t0.158153\t0.157659\t0.158043\t42:09\n2\t0.150255\t0.149575\t0.149918\t43:42\n3\t0.148586\t0.145782\t0.146147\t39:25\n4\t0.145230\t0.144940\t0.145312\t32:42\n5\t0.144837\t0.142363\t0.142794\t32:49\n6\t0.142614\t0.140910\t0.141381\t32:50\n7\t0.140282\t0.139727\t0.140223\t33:27\n8\t0.139156\t0.137941\t0.138437\t43:19\n9\t0.138992\t0.138134\t0.138624\t47:52\n10\t0.136147\t0.136390\t0.136903\t34:05\n11\t0.135508\t0.135124\t0.135669\t39:03\n12\t0.135371\t0.134753\t0.135330\t44:09\n13\t0.134376\t0.133922\t0.134506\t40:50\n14\t0.133351\t0.133514\t0.134096\t45:03\n15\t0.132464\t0.133308\t0.133924\t46:58\n16\t0.131746\t0.132677\t0.133305\t34:27\n17\t0.129510\t0.131462\t0.132088\t37:34\n18\t0.129972\t0.130957\t0.131607\t36:03\n19\t0.128496\t0.130245\t0.130901\t37:28\n20\t0.127367\t0.129550\t0.130230\t45:12\n21\t0.126942\t0.129342\t0.130026\t47:46\n22\t0.125533\t0.128605\t0.129306\t40:11\n23\t0.122462\t0.128378\t0.129091\t40:06\n24\t0.124963\t0.127970\t0.128689\t45:18\n25\t0.124284\t0.127645\t0.128375\t41:10\n26\t0.120596\t0.127453\t0.128191\t46:16\n27\t0.121809\t0.127532\t0.128275\t47:44\n28\t0.121237\t0.127337\t0.128086\t42:42\n29\t0.118522\t0.127245\t0.127996\t42:48\n30\t0.120510\t0.127199\t0.127951\t42:17\n31\t0.119170\t0.127224\t0.127977\t44:47'''","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:38:49.141013Z","iopub.execute_input":"2023-12-08T10:38:49.141325Z","iopub.status.idle":"2023-12-08T10:38:49.145554Z","shell.execute_reply.started":"2023-12-08T10:38:49.141301Z","shell.execute_reply":"2023-12-08T10:38:49.144722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bpp_0 = '''0\t0.171605\t0.172282\t0.172563\t41:28\n1\t0.153931\t0.153551\t0.153885\t26:25\n2\t0.148254\t0.148344\t0.148700\t25:08\n3\t0.146636\t0.144087\t0.144471\t29:03\n4\t0.143474\t0.142159\t0.142586\t26:33\n5\t0.143262\t0.140238\t0.140695\t33:28\n6\t0.141327\t0.139111\t0.139584\t37:35\n7\t0.139394\t0.138686\t0.139189\t35:36\n8\t0.138412\t0.137244\t0.137752\t35:08\n9\t0.138565\t0.136842\t0.137357\t37:34\n10\t0.135920\t0.135916\t0.136450\t34:39\n11\t0.135554\t0.134869\t0.135408\t35:06\n12\t0.135417\t0.134754\t0.135311\t34:16\n13\t0.134570\t0.134145\t0.134720\t34:09\n14\t0.133786\t0.133327\t0.133891\t35:44\n15\t0.132977\t0.133305\t0.133889\t35:42\n16\t0.132718\t0.132165\t0.132770\t35:55\n17\t0.130308\t0.131749\t0.132359\t35:08\n18\t0.131249\t0.131064\t0.131692\t36:09\n19\t0.129823\t0.130766\t0.131393\t35:02\n20\t0.129159\t0.130193\t0.130850\t34:21\n21\t0.128911\t0.129608\t0.130269\t35:00\n22\t0.127578\t0.129145\t0.129818\t35:29\n23\t0.124789\t0.128908\t0.129591\t36:26\n24\t0.127449\t0.128632\t0.129315\t36:02\n25\t0.127041\t0.128150\t0.128848\t36:44\n26\t0.123252\t0.128029\t0.128736\t38:20\n27\t0.125200\t0.127903\t0.128616\t41:13\n28\t0.124595\t0.127716\t0.128431\t37:53\n29\t0.121965\t0.127605\t0.128324\t36:50\n30\t0.124100\t0.127647\t0.128366\t36:20\n31\t0.122755\t0.127645\t0.128365\t37:47'''","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:38:52.053933Z","iopub.execute_input":"2023-12-08T10:38:52.054218Z","iopub.status.idle":"2023-12-08T10:38:52.058669Z","shell.execute_reply.started":"2023-12-08T10:38:52.054197Z","shell.execute_reply":"2023-12-08T10:38:52.057841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport matplotlib.pyplot as plt\ndef get_np_from_history_string(string):\n    train_loss = []\n    for row in string.split('\\n'):\n        train_loss += [np.array(row.split('\\t'))[:4].astype('float')[1:]]\n\n    return np.array(train_loss)\n\nplt.plot(get_np_from_history_string(bpp_0)[:,2])\nplt.plot(get_np_from_history_string(bpp_24_0)[:,2])\nplt.plot(get_np_from_history_string(bpp_24_1)[:,2])\nplt.plot(get_np_from_history_string(bpp_24_2)[:,2])\nplt.plot(get_np_from_history_string(bpp_24_3)[:,2])","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:39:04.179688Z","iopub.execute_input":"2023-12-08T10:39:04.179991Z","iopub.status.idle":"2023-12-08T10:39:04.386093Z","shell.execute_reply.started":"2023-12-08T10:39:04.179970Z","shell.execute_reply":"2023-12-08T10:39:04.385287Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(get_np_from_history_string(bpp_0)[:,0])\nplt.plot(get_np_from_history_string(bpp_24_0)[:,0])\nplt.plot(get_np_from_history_string(bpp_24_1)[:,0])\nplt.plot(get_np_from_history_string(bpp_24_2)[:,0])\nplt.plot(get_np_from_history_string(bpp_24_3)[:,0])","metadata":{"execution":{"iopub.status.busy":"2023-12-08T10:39:08.431188Z","iopub.execute_input":"2023-12-08T10:39:08.431496Z","iopub.status.idle":"2023-12-08T10:39:08.613330Z","shell.execute_reply.started":"2023-12-08T10:39:08.431473Z","shell.execute_reply":"2023-12-08T10:39:08.612650Z"},"trusted":true},"execution_count":null,"outputs":[]}]}