{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\nimport gc\nimport torch\nimport pickle\nimport codecs\nimport gensim\nimport numpy as np\nimport pandas as pd\nimport pickle as pkl\nimport torch.nn as nn\nfrom tqdm import tqdm\nimport seaborn as sns\nimport torch.nn as nn\nimport lightgbm as lgb\nfrom scipy import sparse\nfrom typing import Tuple\nimport torch.nn.functional as F\nfrom sklearn.metrics import log_loss\nfrom text_unidecode import unidecode\nfrom typing import Dict, List, Tuple\nfrom transformers import AutoTokenizer\nfrom sklearn.preprocessing import OneHotEncoder\nfrom torch.utils.data import Dataset, DataLoader\nfrom sklearn.model_selection import StratifiedGroupKFold\nfrom transformers import AutoModel, AutoTokenizer, AutoConfig\nimport warnings; warnings.simplefilter('ignore')\n\ndevice = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\n","metadata":{"pycharm":{"name":"#%%\n"},"execution":{"iopub.status.busy":"2022-07-24T14:01:29.558655Z","iopub.execute_input":"2022-07-24T14:01:29.559531Z","iopub.status.idle":"2022-07-24T14:01:43.320295Z","shell.execute_reply.started":"2022-07-24T14:01:29.559401Z","shell.execute_reply":"2022-07-24T14:01:43.319003Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Huge Ensemble\n#### DeBERTa-Base + DeBERTa-Large + RoBERTa-Large + LightGBM(GoogleNews-W2V)\nIn this notebook, we use 4 different models to generate predictions.\n\n**DeBERTa-Base:**\n- Training [notebook](https://www.kaggle.com/code/yujikomi/train-custommodel) by [YUJI.K](https://www.kaggle.com/yujikomi)\n- Concatenate Discourse Text, [SEP] and Essay Text\n- Pass through Deberta-base-v3 model\n- Apply Mean Pooling on the final hidden states\n- Classify using CrossEntropyLoss into 3 classes\n- Includes Dynamic Padding\n\n**DeBERTa-Large:**\n- Training [notebook](https://www.kaggle.com/code/brandonhu0215/feedback-deberta-large-lb0-619) by [DSML](https://www.kaggle.com/brandonhu0215)\n- 512 max length\n- WeightedLayerPooling (Slightly improve deberta-base model from simple [CLS] head)\n- GroupFold\n- Different learning rates across layers\n- Preprocessing (encoding-resolve+normalize)\n\n**RoBERTa-Large:**\n- Training [notebook](https://www.kaggle.com/code/thedevastator/feedback-roberta-large-training) by [The Devastator](https://www.kaggle.com/thedevastator)\n- 512 max len\n\n**LightGBM(GoogleNews-W2V):**\n- Training [notebook](https://www.kaggle.com/code/mujrush/feedback2-word2vec-lightgbm/notebook) by [MUJ!RUSH!](https://www.kaggle.com/mujrush)\n- LightGBM model using Word2vec.\n- Word2vec represents words in 300 dimensions. By averaging the 300-dimensional vectors of the words in the sentence, the sentence was represented in 300 dimensions.\n\n\n<center>\n    <img src=\"https://i.ibb.co/ygz26m4/power-rangers.png\" style=\"max-height: 600px; border-radius:20px; border: 1px solid;\">    \n<center>","metadata":{"execution":{"iopub.status.busy":"2022-07-23T16:30:09.049310Z","iopub.execute_input":"2022-07-23T16:30:09.050737Z","iopub.status.idle":"2022-07-23T16:30:09.056756Z","shell.execute_reply.started":"2022-07-23T16:30:09.050656Z","shell.execute_reply":"2022-07-23T16:30:09.055631Z"},"pycharm":{"name":"#%% md\n"},"trusted":true}},{"cell_type":"markdown","source":"# Ensemble Configuration","metadata":{}},{"cell_type":"code","source":"WEIGHTS = [0.20, 0.65, 0.05, 0.1]\nMODEL_NAMES = ['deberta', 'deberta_large', 'roberta', 'lgbm']","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# DeBERTa Models\n_____","metadata":{}},{"cell_type":"markdown","source":"### Wait! what is so special about DeBERTa? 🤔\n\n### DeBERTa V1\n\n> [Paper](https://arxiv.org/pdf/2006.03654.pdf)\n> [Official code](https://github.com/microsoft/DeBERTa)\n\nDeBERTa: Decoding-enhanced BERT with disentangled attention: \n\nDeBERTa is a transformer-based neural language model that is (kind of) an improvement to RoBERTa.\nIt improves on BERT and RoBERTa models using two novel techniques. \n\n\n##### Disentangled self-attention mechanism\n\n**The hidden issue of the transformer architecture**\n\nTransformers process sequences assets, this makes them **permutation invariant**.\nIn simple terms: You can shuffle the tokens in the input and the transformer architecture will treat them the same way.\n\n> **Proof that language is not permutation invariant:** \n> \"I am a smart person\" != \"am I a smart person\" (?)\n\n![](https://i.ibb.co/YpFtJwq/0-7fsg-Ie1-Hgw-Pwgf-SD.png)\n\nTo solve this, we use positional encodings: We just add a sin/cos to the input so the transformer will know where exactly in the sequence the token is from.\nBut this creates another problem: Now the transformer is sensitive to the length of the input (and a bit to translation but leave this aside). \n\nA solution to this issue was proposed in the DeBERTa paper.\n\nDeBERTa addresses this by using two learned vectors, which encode content and position, respectively.\n\n##### The Enhanced Mask Decoder\n\nThe second novel technique is the Enhanced Mask Decoder, it incorporates absolute positions [of the sentence] in the decoding layer to predict the masked tokens in model pretraining.\n\n![](https://i.ibb.co/TYTnFcP/0-x-V2-GV7-YX5u4ic47.png)\n\n### DeBERTaV3\n\n> [paper](https://arxiv.org/abs/2111.09543?context=cs)\n> [Official code](https://github.com/microsoft/DeBERTa)\n\n\n**In short:**\n\n- Combine DeBERTa with ELECTRA-style training. \n- Employ a gradient-disentangled embedding sharing as one of the model's building blocks to avoid “tug-of-war” issues.\n\n\n##### Electra's training\n\nThe authors replace the masked language modeling (MLM) with a more sample-efficient pretraining task: **replaced token detection (RTD)**, where the model is trained as a discriminator to **predict whether a token in the input had been corrupted**.\nFor replacing the token in the input, Electra trains a generator that creates adversarial noise that is supposed to \"directly\" be the input on which the discriminator needs to train on.\n\n##### Gradient-disentangled embedding sharing (GDES)\n\nIn ELECTRA, the discriminator and the generator share the same token embeddings. This mechanism can however hurt training efficiency, as the training losses of the discriminator and the generator tend to pull token embeddings in different directions. \nThe simple solution: The generator shares its embeddings with the discriminator but stops the gradients in the discriminator from backpropagating to the generator embeddings.\n","metadata":{}},{"cell_type":"markdown","source":"## Model 1: Deberta-Base","metadata":{}},{"cell_type":"markdown","source":"#### Configurations","metadata":{}},{"cell_type":"code","source":"INPUT_DIR = '../input/feedback-prize-effectiveness/'\n\nclass CFG:\n    CVs = []\n    seed = 42\n    lr = 3e-5\n    epochs = 3\n    n_fold = 5\n    apex = True\n    fast = True\n    AMP = False\n    n_splits = 5\n    train = True\n    wandb = False\n    max_len = 512\n    dropout = 0.1\n    min_lr = 1e-6\n    batch_size = 8\n    freezing = True\n    print_freq = 50\n    target_size = 3\n    num_workers = 0\n    num_cycles = 0.5\n    n_accumulate = 1\n    scheduler = 'cosine'\n    weigth_decay = 0.01\n    num_warmup_steps = 0\n    trn_fold = [0, 1, 2, 3, 4]\n    gradient_checkpointing = True\n    model = '../input/deberta-v3-base/deberta-v3-base'","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:43.323226Z","iopub.execute_input":"2022-07-24T14:01:43.323690Z","iopub.status.idle":"2022-07-24T14:01:43.333866Z","shell.execute_reply.started":"2022-07-24T14:01:43.323645Z","shell.execute_reply":"2022-07-24T14:01:43.332543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Helper Function","metadata":{}},{"cell_type":"code","source":"def criterion(outputs, labels):\n    return nn.CrossEntropyLoss()(outputs, labels)\n\ndef softmax(z):\n    assert len(z.shape) == 2\n    s = np.max(z, axis=1)\n    s = s[:, np.newaxis]\n    e_x = np.exp(z - s)\n    div = np.sum(e_x, axis=1)\n    div = div[:, np.newaxis]\n    return e_x / div\n\ndef freeze(module):\n    for parameter in module.parameters():\n        parameter.requires_grad = False\n        \ndef get_freezed_parameters(module):\n    freezed_parameters = []\n    for name, parameter in module.named_parameters():\n        if not parameter.requires_grad:\n            freezed_parameters.append(name)\n    return freezed_parameters\n\ndef get_essay(essay_id, is_train=True):\n    parent_path = INPUT_DIR + 'train' if is_train else INPUT_DIR + 'test'\n    essay_path = os.path.join(parent_path, f\"{essay_id}.txt\")\n    essay_text = open(essay_path, 'r').read()\n    return essay_text","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:43.336148Z","iopub.execute_input":"2022-07-24T14:01:43.337097Z","iopub.status.idle":"2022-07-24T14:01:43.351648Z","shell.execute_reply.started":"2022-07-24T14:01:43.337053Z","shell.execute_reply":"2022-07-24T14:01:43.350352Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Preprocessing","metadata":{}},{"cell_type":"code","source":"# Testing Data\ntest = pd.read_csv(INPUT_DIR + 'test.csv')\ntest['essay_text'] = test['essay_id'].apply(lambda x: get_essay(x, is_train=False))","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:43.357199Z","iopub.execute_input":"2022-07-24T14:01:43.357780Z","iopub.status.idle":"2022-07-24T14:01:43.398670Z","shell.execute_reply.started":"2022-07-24T14:01:43.357750Z","shell.execute_reply":"2022-07-24T14:01:43.397399Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if CFG.fast: tokenizer = AutoTokenizer.from_pretrained(CFG.model, use_fast=True)\nelse: tokenizer = AutoTokenizer.from_pretrained(CFG.model)\nCFG.tokenizer = tokenizer","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:43.400780Z","iopub.execute_input":"2022-07-24T14:01:43.401285Z","iopub.status.idle":"2022-07-24T14:01:44.235661Z","shell.execute_reply.started":"2022-07-24T14:01:43.401239Z","shell.execute_reply":"2022-07-24T14:01:44.234352Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Normalization","metadata":{}},{"cell_type":"code","source":"def replace_encoding_with_utf8(error: UnicodeError) -> Tuple[bytes, int]: return error.object[error.start : error.end].encode(\"utf-8\"), error.end\ndef replace_decoding_with_cp1252(error: UnicodeError) -> Tuple[str, int]: return error.object[error.start : error.end].decode(\"cp1252\"), error.end\ncodecs.register_error(\"replace_encoding_with_utf8\", replace_encoding_with_utf8)\ncodecs.register_error(\"replace_decoding_with_cp1252\", replace_decoding_with_cp1252)\n\ndef resolve_encodings_and_normalize(text: str) -> str:\n    text = (text.encode(\"raw_unicode_escape\").decode(\"utf-8\", errors = \"replace_decoding_with_cp1252\").encode(\"cp1252\", errors = \"replace_encoding_with_utf8\").decode(\"utf-8\", errors = \"replace_decoding_with_cp1252\"))\n    text = unidecode(text)\n    return text","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:44.237830Z","iopub.execute_input":"2022-07-24T14:01:44.239719Z","iopub.status.idle":"2022-07-24T14:01:44.252872Z","shell.execute_reply.started":"2022-07-24T14:01:44.239615Z","shell.execute_reply":"2022-07-24T14:01:44.251247Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test['discourse_text'] = test['discourse_text'].apply(lambda x : resolve_encodings_and_normalize(x))\ntest['essay_text'] = test['essay_text'].apply(lambda x : resolve_encodings_and_normalize(x))\ntest['text'] = test['discourse_type'] + ' ' + test['discourse_text'] + '[SEP]' + test['essay_text']\ntest['label'] = np.nan","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:44.254481Z","iopub.execute_input":"2022-07-24T14:01:44.255041Z","iopub.status.idle":"2022-07-24T14:01:44.277399Z","shell.execute_reply.started":"2022-07-24T14:01:44.254997Z","shell.execute_reply":"2022-07-24T14:01:44.275862Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Dataset + Dynamic padding","metadata":{}},{"cell_type":"code","source":"class TestDataset(Dataset):\n    def __init__(self, cfg, df):\n        self.cfg = cfg\n        self.text = df['text'].values\n    def __len__(self): return len(self.text)\n    def __getitem__(self, item):\n        inputs = self.cfg.tokenizer.encode_plus(self.text[item], truncation = True, add_special_tokens = True, max_length = self.cfg.max_len)\n        samples = {'input_ids': inputs['input_ids'], 'attention_mask': inputs['attention_mask'], }\n        if 'token_type_ids' in inputs: samples['token_type_ids'] = inputs['token_type_ids']\n        return samples\n\nclass Collate:\n    def __init__(self, tokenizer, isTrain=True):\n        self.isTrain = isTrain\n        self.tokenizer = tokenizer\n\n    def __call__(self, batch):\n        output = dict()\n        output[\"input_ids\"] = [sample[\"input_ids\"] for sample in batch]\n        output[\"attention_mask\"] = [sample[\"attention_mask\"] for sample in batch]\n        if self.isTrain: output[\"target\"] = [sample[\"target\"] for sample in batch]\n        batch_max = max([len(ids) for ids in output[\"input_ids\"]])\n        if self.tokenizer.padding_side == \"right\":\n            output[\"input_ids\"] = [s + (batch_max - len(s)) * [self.tokenizer.pad_token_id] for s in output[\"input_ids\"]]\n            output[\"attention_mask\"] = [s + (batch_max - len(s)) * [0] for s in output[\"attention_mask\"]]\n        else:\n            output[\"input_ids\"] = [(batch_max - len(s)) * [self.tokenizer.pad_token_id] + s for s in output[\"input_ids\"]]\n            output[\"attention_mask\"] = [(batch_max - len(s)) * [0] + s for s in output[\"attention_mask\"]]\n        output[\"input_ids\"] = torch.tensor(output[\"input_ids\"], dtype=torch.long)\n        output[\"attention_mask\"] = torch.tensor(output[\"attention_mask\"], dtype=torch.long)\n        if self.isTrain: output[\"target\"] = torch.tensor(output[\"target\"], dtype=torch.long)\n        return output","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:44.280048Z","iopub.execute_input":"2022-07-24T14:01:44.281082Z","iopub.status.idle":"2022-07-24T14:01:44.300758Z","shell.execute_reply.started":"2022-07-24T14:01:44.281028Z","shell.execute_reply":"2022-07-24T14:01:44.299167Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### The Model","metadata":{}},{"cell_type":"code","source":"class MeanPooling(nn.Module):\n    def __init__(self):\n        super(MeanPooling, self).__init__()\n        \n    def forward(self, last_hidden_state, attention_mask):\n        input_mask_expanded = attention_mask.unsqueeze(-1).expand(last_hidden_state.size()).float()\n        sum_embeddings = torch.sum(last_hidden_state * input_mask_expanded, 1)\n        sum_mask = input_mask_expanded.sum(1)\n        sum_mask = torch.clamp(sum_mask, min=1e-9) #\n        mean_embeddings = sum_embeddings / sum_mask\n        return mean_embeddings\n\ndef inference_fn(test_loader, model, device):\n    preds = []\n    model.eval()\n    model.to(device)\n    tk0 = tqdm(test_loader, total=len(test_loader))\n    for data in tk0:\n        ids = data['input_ids'].to(device, dtype = torch.long)\n        mask = data['attention_mask'].to(device, dtype = torch.long)\n        with torch.no_grad():\n            y_preds = model(ids, mask)\n        y_preds = softmax(y_preds.to('cpu').numpy())\n        preds.append(y_preds)\n    predictions = np.concatenate(preds)\n    return predictions","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:44.302809Z","iopub.execute_input":"2022-07-24T14:01:44.303899Z","iopub.status.idle":"2022-07-24T14:01:44.318650Z","shell.execute_reply.started":"2022-07-24T14:01:44.303849Z","shell.execute_reply":"2022-07-24T14:01:44.317371Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class FeedBackModel(nn.Module):\n    def __init__(self, model_name):\n        super(FeedBackModel, self).__init__()\n        self.model = AutoModel.from_pretrained(model_name)\n        if CFG.gradient_checkpointing: (self.model).gradient_checkpointing_enable()\n        if CFG.freezing:\n            freeze((self.model).embeddings)\n            freeze((self.model).encoder.layer[:2])\n            CFG.after_freezed_parameters = filter(lambda parameter: parameter.requires_grad, (self.model).parameters())\n        self.config = AutoConfig.from_pretrained(model_name)\n        self.drop = nn.Dropout(p=CFG.dropout)\n        self.pooler = MeanPooling()\n        self.fc = nn.Linear(self.config.hidden_size, CFG.target_size)\n        \n    def forward(self, ids, mask):\n        out = self.model(input_ids = ids, attention_mask = mask, output_hidden_states = False)\n        out = self.pooler(out.last_hidden_state, mask)\n        out = self.drop(out)\n        outputs = self.fc(out)\n        return outputs","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:44.326023Z","iopub.execute_input":"2022-07-24T14:01:44.328734Z","iopub.status.idle":"2022-07-24T14:01:44.340366Z","shell.execute_reply.started":"2022-07-24T14:01:44.328671Z","shell.execute_reply":"2022-07-24T14:01:44.339149Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Deberta-Base Inference","metadata":{}},{"cell_type":"code","source":"testDataset = TestDataset(CFG, test)\ntest_loader = DataLoader(\n                          testDataset,\n                          shuffle = False,\n                          drop_last = False,\n                          pin_memory = True,\n                          batch_size = CFG.batch_size,\n                          num_workers = CFG.num_workers,\n                          collate_fn = Collate(CFG.tokenizer, isTrain = False)\n                        )\n\ndeberta_predictions = []\nfor i in CFG.trn_fold:\n    model = FeedBackModel(CFG.model)\n    model.load_state_dict(torch.load('../input/dbv3basemodels202279/models-deberta-v3-base-deberta-v3-base_fold' + str(i) +'_best.pth'))\n    prediction = inference_fn(test_loader, model, device)\n    deberta_predictions.append(prediction)\n    torch.cuda.empty_cache()\n    gc.collect()\n\ndeb_adequate = []\ndeb_effective = []\ndeb_ineffective = []\n\nfor x in deberta_predictions:\n    deb_ineffective.append(x[:, 0])\n    deb_adequate.append(x[:, 1])\n    deb_effective.append(x[:, 2])\n\ndeb_ineffective = pd.DataFrame(deb_ineffective).T\ndeb_adequate = pd.DataFrame(deb_adequate).T\ndeb_effective = pd.DataFrame(deb_effective).T","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:01:44.342311Z","iopub.execute_input":"2022-07-24T14:01:44.343148Z","iopub.status.idle":"2022-07-24T14:02:48.475984Z","shell.execute_reply.started":"2022-07-24T14:01:44.343104Z","shell.execute_reply":"2022-07-24T14:02:48.474381Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model 2: Deberta-Large","metadata":{}},{"cell_type":"markdown","source":"#### Configurations","metadata":{}},{"cell_type":"code","source":"class CFG:\n    seed = 42\n    n_fold = 4\n    max_len = 512\n    batch_size = 32\n    num_workers = 2\n    model = \"microsoft/deberta-large\"\n    path = \"../input/feedback-deberta-large-051/\"\n    config_path = \"../input/feedback-deberta-large-051/\" + 'config.pth'","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:05:48.814077Z","iopub.execute_input":"2022-07-24T14:05:48.814523Z","iopub.status.idle":"2022-07-24T14:05:48.908944Z","shell.execute_reply.started":"2022-07-24T14:05:48.814491Z","shell.execute_reply":"2022-07-24T14:05:48.907537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Helper Functions","metadata":{}},{"cell_type":"code","source":"def replace_encoding_with_utf8(error: UnicodeError) -> Tuple[bytes, int]: return error.object[error.start : error.end].encode(\"utf-8\"), error.end\ndef replace_decoding_with_cp1252(error: UnicodeError) -> Tuple[str, int]: return error.object[error.start : error.end].decode(\"cp1252\"), error.end\ncodecs.register_error(\"replace_encoding_with_utf8\", replace_encoding_with_utf8)\ncodecs.register_error(\"replace_decoding_with_cp1252\", replace_decoding_with_cp1252)\n\ndef resolve_encodings_and_normalize(text: str) -> str:\n    text = (text.encode(\"raw_unicode_escape\").decode(\"utf-8\", errors = \"replace_decoding_with_cp1252\").encode(\"cp1252\", errors = \"replace_encoding_with_utf8\").decode(\"utf-8\", errors = \"replace_decoding_with_cp1252\"))\n    text = unidecode(text)\n    return text\n\ndef fetch_essay(essay_id: str, txt_dir: str):\n    essay_path = os.path.join(COMP_DIR + txt_dir, essay_id + '.txt')\n    essay_text = open(essay_path, 'r').read()\n    return essay_text\n\ndef prepare_input(cfg, text, text_2=None):\n    inputs = cfg.tokenizer(text, text_2, padding = \"max_length\", add_special_tokens = True, max_length = cfg.max_len, truncation = True)\n    for k, v in inputs.items(): inputs[k] = torch.tensor(v, dtype=torch.long)\n    return inputs\n\ndef inference_fn(test_loader, model, device):\n    preds = []\n    model.eval()\n    model.to(device)\n    tk0 = tqdm(test_loader, total=len(test_loader))\n    for inputs in tk0:\n        for k, v in inputs.items():\n            inputs[k] = v.to(device)\n        with torch.no_grad():\n            output = model(inputs)\n        preds.append(F.softmax(output).to('cpu').numpy())\n    return np.concatenate(preds)\n\ndef show_gradient(df, n_row=None):\n    if not n_row: n_row = 5\n    return df.head(n_row).assign(all_mean=lambda x: x.mean(axis=1)).style.background_gradient(cmap=cm, axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:02:48.478028Z","iopub.execute_input":"2022-07-24T14:02:48.478456Z","iopub.status.idle":"2022-07-24T14:02:48.497369Z","shell.execute_reply.started":"2022-07-24T14:02:48.478414Z","shell.execute_reply":"2022-07-24T14:02:48.495783Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Data Loading","metadata":{}},{"cell_type":"code","source":"N_ROW = 10\n\npd.set_option('display.precision', 4)\ncm = sns.light_palette('green', as_cmap=True)\nprops_param = \"color:white; font-weight:bold; background-color:green;\"\nCOMP_DIR = \"../input/feedback-prize-effectiveness/\"\nDEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')\ntest_path = COMP_DIR + \"test.csv\"\nsubmission_path = COMP_DIR + \"sample_submission.csv\"\ntest_origin = pd.read_csv(test_path)\nsubmission_origin = pd.read_csv(submission_path)\ndata_path = \"../input/feedback-prize-effectiveness/train.csv\"\ncols_list = ['essay_id', 'discourse_text']\nidxs_list = [49, 80, 945, 947, 1870]\ntemp = pd.read_csv(data_path, usecols=cols_list).loc[idxs_list, :]","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:02:48.499609Z","iopub.execute_input":"2022-07-24T14:02:48.500927Z","iopub.status.idle":"2022-07-24T14:02:48.785998Z","shell.execute_reply.started":"2022-07-24T14:02:48.500880Z","shell.execute_reply":"2022-07-24T14:02:48.784767Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp['discourse_text_UPD'] = temp['discourse_text'].apply(resolve_encodings_and_normalize)\ntemp['essay_text'] = temp['essay_id'].transform(fetch_essay, txt_dir='train')\ntemp['essay_text_UPD'] = temp['essay_text'].apply(resolve_encodings_and_normalize)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:02:48.787490Z","iopub.execute_input":"2022-07-24T14:02:48.787921Z","iopub.status.idle":"2022-07-24T14:02:48.820397Z","shell.execute_reply.started":"2022-07-24T14:02:48.787878Z","shell.execute_reply":"2022-07-24T14:02:48.819028Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for n, row in enumerate(temp.iterrows()):\n    indx, data = row\n    disc_text = data.discourse_text\n    disc_text_upd = data.discourse_text_UPD\n    print(f'\\nN{n} === index: {indx} ===')\n    print(f'\\n>>> origin text:')\n    print(repr(disc_text))\n    print(f'\\n>>> updated text:')\n    print(repr(disc_text_upd))","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:02:48.822050Z","iopub.execute_input":"2022-07-24T14:02:48.823383Z","iopub.status.idle":"2022-07-24T14:02:48.835989Z","shell.execute_reply.started":"2022-07-24T14:02:48.823289Z","shell.execute_reply":"2022-07-24T14:02:48.834237Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Deberta Large - The Model","metadata":{}},{"cell_type":"code","source":"class TestDataset(Dataset):\n    def __init__(self, cfg, df):\n        self.cfg = cfg\n        self.text = df['text'].values\n    def __len__(self): return len(self.text)\n    def __getitem__(self, item):\n        text = self.text[item]\n        inputs = prepare_input(self.cfg, text)\n        return inputs\n\nclass CustomModel(nn.Module):\n    def __init__(self, cfg, config_path=None, pretrained=False):\n        super().__init__()\n        self.cfg = cfg\n        if config_path is None: self.config = AutoConfig.from_pretrained(cfg.model, output_hidden_states=True)\n        else: self.config = torch.load(config_path)\n        if pretrained: self.model = AutoModel.from_pretrained(cfg.model, config=self.config)\n        else: self.model = AutoModel.from_config(self.config)\n        self.bilstm = nn.LSTM(self.config.hidden_size, (self.config.hidden_size) // 2, num_layers=2, dropout=self.config.hidden_dropout_prob, batch_first=True, bidirectional=True)\n        self.dropout1 = nn.Dropout(0.1)\n        self.dropout2 = nn.Dropout(0.2)\n        self.dropout3 = nn.Dropout(0.3)\n        self.dropout4 = nn.Dropout(0.4)\n        self.dropout5 = nn.Dropout(0.5)\n        self.output = nn.Sequential( nn.Linear(self.config.hidden_size, 3) )\n                \n    def _init_weights(self, module):\n        if isinstance(module, nn.Linear):\n            module.weight.data.normal_(mean=0.0, std=self.config.initializer_range)\n            if module.bias is not None:\n                module.bias.data.zero_()\n        elif isinstance(module, nn.Embedding):\n            module.weight.data.normal_(mean=0.0, std=self.config.initializer_range)\n            if module.padding_idx is not None:\n                module.weight.data[module.padding_idx].zero_()\n        elif isinstance(module, nn.LayerNorm):\n            module.bias.data.zero_()\n            module.weight.data.fill_(1.0)\n\n    def forward(self, inputs):\n        sequence_output = self.model(**inputs)[0][:, 0, :]\n        logits1 = self.output(self.dropout1(sequence_output))\n        logits2 = self.output(self.dropout2(sequence_output))\n        logits3 = self.output(self.dropout3(sequence_output))\n        logits4 = self.output(self.dropout4(sequence_output))\n        logits5 = self.output(self.dropout5(sequence_output))\n        logits = (logits1 + logits2 + logits3 + logits4 + logits5) / 5\n        return logits","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:02:48.838421Z","iopub.execute_input":"2022-07-24T14:02:48.838963Z","iopub.status.idle":"2022-07-24T14:02:48.861771Z","shell.execute_reply.started":"2022-07-24T14:02:48.838917Z","shell.execute_reply":"2022-07-24T14:02:48.860163Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CFG.tokenizer = AutoTokenizer.from_pretrained(CFG.path + 'tokenizer')\n\ndf = test_origin.copy()\nSEP = CFG.tokenizer.sep_token\ndf['discourse_text'] = df['discourse_text'].apply(resolve_encodings_and_normalize)\ndf['essay_text'] = df['essay_id'].transform(fetch_essay, txt_dir='test')\ndf['essay_text'] = df['essay_text'].apply(resolve_encodings_and_normalize)\ndf['text'] = df['discourse_type'] + ' ' + df['discourse_text'] + SEP + df['essay_text']","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:05:48.911844Z","iopub.execute_input":"2022-07-24T14:05:48.912520Z","iopub.status.idle":"2022-07-24T14:05:48.939297Z","shell.execute_reply.started":"2022-07-24T14:05:48.912457Z","shell.execute_reply":"2022-07-24T14:05:48.937965Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_dataset = TestDataset(CFG, df)\ntest_loader = DataLoader(test_dataset, batch_size = CFG.batch_size, shuffle = False, num_workers = CFG.num_workers, pin_memory = True, drop_last = False)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:05:48.941821Z","iopub.execute_input":"2022-07-24T14:05:48.942681Z","iopub.status.idle":"2022-07-24T14:05:48.950610Z","shell.execute_reply.started":"2022-07-24T14:05:48.942637Z","shell.execute_reply":"2022-07-24T14:05:48.949051Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Deberta-Large Inference","metadata":{}},{"cell_type":"code","source":"deberta_large_predictions = []\nfor fold in range(CFG.n_fold):\n    model = CustomModel(CFG, config_path=CFG.config_path, pretrained=False)\n    state = torch.load(CFG.path + f\"{CFG.model.replace('/', '-')}_fold{fold}_best.pth\", map_location=torch.device('cpu'))\n    model.load_state_dict(state['model'])\n    prediction = inference_fn(test_loader, model, DEVICE)\n    deberta_large_predictions.append(prediction)\n    del model, state, prediction; gc.collect()\n    torch.cuda.empty_cache()\n\ndeb_large_adequate = []\ndeb_large_effective = []\ndeb_large_ineffective = []\n\nfor x in deberta_large_predictions:\n    deb_large_ineffective.append(x[:, 0])\n    deb_large_adequate.append(x[:, 1])\n    deb_large_effective.append(x[:, 2])\n\ndeb_large_ineffective = pd.DataFrame(deb_large_ineffective).T\ndeb_large_adequate = pd.DataFrame(deb_large_adequate).T\ndeb_large_effective = pd.DataFrame(deb_large_effective).T","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:05:48.952802Z","iopub.execute_input":"2022-07-24T14:05:48.953351Z","iopub.status.idle":"2022-07-24T14:07:40.399048Z","shell.execute_reply.started":"2022-07-24T14:05:48.953271Z","shell.execute_reply":"2022-07-24T14:07:40.397458Z"},"_kg_hide-input":false,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ReBERTa Models\n_____","metadata":{}},{"cell_type":"markdown","source":"Couple of years ago when the a paper was published with the claim: **The original BERT architecture can outperform any other model released after it**.\n\nShocking. right? \n\nThe authors managed to do it using some heavy hyperparameters search and many other training scheme tricks. They simply named their method **R**obustly **O**ptimized **BERT** pretraining **A**pproach or simply **RoBERTa**.\nWhich to this day remain as one of the top high performing transformers available.\n\n\n**In short: RoBERTa's Hyperparameters change from BERT**\n\n- Longer training time.\n- Larger training data (x10, from 16G to 160GB).\n- Larger batch size (from 256 to 8k).\n- The removal of the NSP task.\n- Bigger vocabulary size (from 30k to 50k).\n- Longer sequences are used as input (but still keep the limitation of 512 tokens).\n- Dynamic masking.","metadata":{}},{"cell_type":"markdown","source":"## Model 3: Roberta-Large","metadata":{}},{"cell_type":"markdown","source":"#### Configurations","metadata":{}},{"cell_type":"code","source":"class CFG:\n    n_fold = 5\n    batch = 16\n    max_len = 512\n    num_workers = 2\n    path = \"../input/robertalarge\"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class TestDataset(Dataset):\n    def __init__(self, cfg, df):\n        self.cfg = cfg\n        self.essay = df['essay'].values\n        self.discourse = df['discourse'].values\n\n    def __len__(self): return len(self.discourse)\n    \n    def __getitem__(self, item):\n        discourse = self.discourse[item]\n        essay = self.essay[item]\n        inputs = prepare_input(self.cfg, discourse, essay)\n        return inputs\n        \nclass FeedBackModel(nn.Module):\n    def __init__(self, model_path):\n        super(FeedBackModel, self).__init__()\n        self.model = AutoModel.from_pretrained(model_path)\n        self.linear = nn.Linear(1024, 3)\n\n    def forward(self, inputs):\n        last_hidden_states = self.model(**inputs)[0][:, 0, :]\n        outputs = self.linear(last_hidden_states)\n        return outputs","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:07:40.403104Z","iopub.execute_input":"2022-07-24T14:07:40.403617Z","iopub.status.idle":"2022-07-24T14:07:40.414258Z","shell.execute_reply.started":"2022-07-24T14:07:40.403567Z","shell.execute_reply":"2022-07-24T14:07:40.412858Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"CFG.tokenizer = AutoTokenizer.from_pretrained(CFG.path)\n\ndf = test_origin.copy()\n\ntxt_sep = \" \"\ndf['discourse'] = df['discourse_type'].str.strip() + txt_sep + df['discourse_text'].str.strip()\ndf['essay'] = df['essay_id'].transform(fetch_essay, txt_dir='test').str.strip()\n\ntest_dataset = TestDataset(CFG, df)\ntest_loader = DataLoader(test_dataset, batch_size=CFG.batch, shuffle=False, num_workers=CFG.num_workers, pin_memory=True, drop_last=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:07:40.415920Z","iopub.execute_input":"2022-07-24T14:07:40.416718Z","iopub.status.idle":"2022-07-24T14:07:40.680060Z","shell.execute_reply.started":"2022-07-24T14:07:40.416674Z","shell.execute_reply":"2022-07-24T14:07:40.678855Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Roberta-Large Inference","metadata":{}},{"cell_type":"code","source":"gc.collect()\nroberta_predicts = []\nfor model_path in os.listdir('../input/feedback-roberta-models'):\n    if 'data_1' in model_path:\n        model = pkl.load(open('../input/feedback-roberta-models/' + model_path, 'rb'))    \n        prediction = inference_fn(test_loader, model, DEVICE)\n        roberta_predicts.append(prediction)\n        del model, prediction\n        torch.cuda.empty_cache()    \n        gc.collect()\ngc.collect()\n\nrob_adequate = []\nrob_effective = []\nrob_ineffective = []\n\nfor x in roberta_predicts:\n    rob_ineffective.append(x[:, 0])\n    rob_adequate.append(x[:, 1])\n    rob_effective.append(x[:, 2])\n\nrob_ineffective = pd.DataFrame(rob_ineffective).T\nrob_adequate = pd.DataFrame(rob_adequate).T\nrob_effective = pd.DataFrame(rob_effective).T","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:07:40.681948Z","iopub.execute_input":"2022-07-24T14:07:40.682419Z","iopub.status.idle":"2022-07-24T14:08:56.725995Z","shell.execute_reply.started":"2022-07-24T14:07:40.682374Z","shell.execute_reply":"2022-07-24T14:08:56.724706Z"},"_kg_hide-input":false,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model 4: LightGBM","metadata":{}},{"cell_type":"markdown","source":"#### General Configurations","metadata":{}},{"cell_type":"code","source":"class CFG:\n    seed = 42\n    n_folds = 4","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:08:56.728151Z","iopub.execute_input":"2022-07-24T14:08:56.728657Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### LightGBM Hyperparameters","metadata":{}},{"cell_type":"code","source":"num_rounds = 1000    \n\nparams = {}\nparams['num_class'] = 3\nparams[\"max_depth\"] = 10\nparams[\"verbosity\"] = -1\nparams[\"num_leaves\"] = 52\nparams['boosting'] = 'gbdt'\nparams[\"bagging_freq\"] = 8\nparams[\"random_state\"] = 42\nparams[\"bagging_seed\"] = 10\nparams[\"lambda_l2\"] = 0.0256\nparams['is_unbalance'] = True\nparams[\"learning_rate\"] = 0.05\nparams[\"min_data_in_leaf\"] = 10\nparams['metric'] = 'multi_logloss'\nparams[\"objective\"] = 'multiclass'\nparams[\"feature_fraction\"] = 0.503\nparams[\"bagging_fraction\"] = 0.741","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:08:56.728151Z","iopub.execute_input":"2022-07-24T14:08:56.728657Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Data Loading","metadata":{}},{"cell_type":"code","source":"INPUT_DIR = \"../input/feedback-prize-effectiveness/\"\n\ndef get_train_essay(essay_id):\n    essay_path = os.path.join(INPUT_DIR,f'train/{essay_id}.txt')\n    essay_text = open(essay_path,'r').read()\n    return essay_text\n\ndef get_test_essay(essay_id):\n    essay_path = os.path.join(INPUT_DIR,f'test/{essay_id}.txt')\n    essay_text = open(essay_path,'r').read()\n    return essay_text\n\ntrain = pd.read_csv(INPUT_DIR+'train.csv')\ntest = pd.read_csv(INPUT_DIR+'test.csv')\ntrain['essay_text'] = train['essay_id'].apply(get_train_essay)\ntest['essay_text'] = test['essay_id'].apply(get_test_essay)\n\ndef set_seed(seed=42):\n    np.random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n\nset_seed(CFG.seed)\n\neffectiveness_map = {'Ineffective':0, 'Adequate':1, 'Effective':2}\ntrain['target'] = train['discourse_effectiveness'].map(effectiveness_map)\n\nfor fold, (_,val_idx) in enumerate(StratifiedGroupKFold(n_splits=CFG.n_folds,shuffle=True,random_state=CFG.seed).split(X=train, y=train['target'], groups=train.essay_id)):\n    train.loc[val_idx,'kfold'] = fold\nword2vec_model = gensim.models.KeyedVectors.load_word2vec_format('../input/google-news/GoogleNews-vectors-negative300.bin', binary=True)\n\ndef avg_feature_vector(sentence, model, num_features):\n    words = sentence.replace('\\n',\" \").replace(',',' ').replace('.',\" \").split()\n    feature_vec = np.zeros((num_features,),dtype=\"float32\")\n    i=0\n    for word in words:\n        try: feature_vec = np.add(feature_vec, model[word])\n        except KeyError as error:\n            feature_vec\n            i = i + 1\n    if len(words) > 0:\n        feature_vec = np.divide(feature_vec, len(words)- i)\n    return feature_vec","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:08:56.728151Z","iopub.execute_input":"2022-07-24T14:08:56.728657Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### LightGBM Training & Inference","metadata":{}},{"cell_type":"code","source":"oof_score = 0\ny_test_pred = np.zeros((test.shape[0], 3))\n\nfor fold in range(CFG.n_folds):\n    print(f'=============fold:{fold}==================')\n    train_fold = train[train['kfold']!=fold].reset_index(drop=True)\n    valid_fold = train[train['kfold']==fold].reset_index(drop=True)\n\n    word2vec_train_disc_text = np.zeros((len(train_fold.index),300),dtype=\"float32\")\n    word2vec_valid_disc_text = np.zeros((len(valid_fold.index),300),dtype=\"float32\")\n    word2vec_test_disc_text = np.zeros((len(test.index),300),dtype=\"float32\")\n    for i in range(len(train_fold.index)): word2vec_train_disc_text[i] = avg_feature_vector(train_fold[\"discourse_text\"][i], word2vec_model, 300)\n    for i in range(len(valid_fold.index)): word2vec_valid_disc_text[i] = avg_feature_vector(valid_fold[\"discourse_text\"][i], word2vec_model, 300)\n    for i in range(len(test.index)): word2vec_test_disc_text[i] = avg_feature_vector(test[\"discourse_text\"][i], word2vec_model, 300)\n\n    word2vec_train_essay_text = np.zeros((len(train_fold.index),300),dtype=\"float32\")\n    word2vec_valid_essay_text = np.zeros((len(valid_fold.index),300),dtype=\"float32\")\n    word2vec_test_essay_text = np.zeros((len(test.index),300),dtype=\"float32\")\n    for i in range(len(train_fold.index)): word2vec_train_essay_text[i] = avg_feature_vector(train_fold[\"essay_text\"][i], word2vec_model, 300)\n    for i in range(len(valid_fold.index)): word2vec_valid_essay_text[i] = avg_feature_vector(valid_fold[\"essay_text\"][i], word2vec_model, 300)\n    for i in range(len(test.index)): word2vec_test_essay_text[i] = avg_feature_vector(test[\"essay_text\"][i], word2vec_model, 300)\n\n    ohe = OneHotEncoder()\n    train_type_ohe = sparse.csr_matrix(ohe.fit_transform(train_fold['discourse_type'].values.reshape(-1,1)))\n    valid_type_ohe = sparse.csr_matrix(ohe.transform(valid_fold['discourse_type'].values.reshape(-1,1)))\n    test_type_ohe = sparse.csr_matrix(ohe.transform(test['discourse_type'].values.reshape(-1,1)))\n\n    Xtrain_word2vec = sparse.hstack((train_type_ohe,word2vec_train_disc_text,word2vec_train_essay_text))\n    Xvalid_word2vec = sparse.hstack((valid_type_ohe,word2vec_valid_disc_text,word2vec_valid_essay_text))\n    test_word2vec = sparse.hstack((test_type_ohe,word2vec_test_disc_text,word2vec_test_essay_text))\n\n    #lgbm\n    lgtrain = lgb.Dataset(Xtrain_word2vec, label=train_fold['target'].ravel())\n    lgvalidation = lgb.Dataset(Xvalid_word2vec, label=valid_fold['target'].ravel())\n\n    model = lgb.train(params, lgtrain, num_rounds, valid_sets=[lgtrain, lgvalidation], early_stopping_rounds=100, verbose_eval=100)\n    y_pred = model.predict(Xvalid_word2vec, num_iteration=model.best_iteration)\n    y_test_pred += model.predict(test_word2vec, num_iteration=model.best_iteration)\n\n    score = log_loss(valid_fold['target'], y_pred)\n    oof_score += score\n\n    print(f'Fold:{fold},valid score:{score}')","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:08:56.728151Z","iopub.execute_input":"2022-07-24T14:08:56.728657Z"},"_kg_hide-input":false,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### LightGBM Predictions Aggregation","metadata":{}},{"cell_type":"code","source":"y_test_pred = y_test_pred / float(CFG.n_folds)\noof_score /= float(CFG.n_folds)\nprint(\"Aggregate OOF Score: {}\".format(oof_score))\n\nlgbm_adequate = y_test_pred[:,1]\nlgbm_effective = y_test_pred[:,2]\nlgbm_ineffective = y_test_pred[:,0]\n\nlgbm_adequate = pd.DataFrame(lgbm_adequate)\nlgbm_effective = pd.DataFrame(lgbm_effective)\nlgbm_ineffective = pd.DataFrame(lgbm_ineffective)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T14:08:56.728151Z","iopub.execute_input":"2022-07-24T14:08:56.728657Z"},"_kg_hide-input":false,"_kg_hide-output":false,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ensembling","metadata":{}},{"cell_type":"code","source":"submission = pd.read_csv('../input/feedback-prize-effectiveness/sample_submission.csv')\n\nineffective_ = pd.concat([deb_ineffective, deb_large_ineffective, rob_ineffective, lgbm_ineffective], keys = MODEL_NAMES, axis = 1)\nadequate_ = pd.concat([deb_adequate, deb_large_adequate, rob_adequate, lgbm_adequate], keys = MODEL_NAMES, axis = 1)\neffective_ = pd.concat([deb_effective, deb_large_effective, rob_effective, lgbm_effective], keys = MODEL_NAMES, axis = 1)","metadata":{"_kg_hide-input":false,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_gradient(ineffective_, N_ROW)\nshow_gradient(adequate_, N_ROW)\nshow_gradient(effective_, N_ROW)","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"w_ = WEIGHTS\nd_ = [('Ineffective', ineffective_), ('Adequate', adequate_), ('Effective', effective_)]\n\nfor x in d_:\n    col_name, df = x\n    submission[col_name] = pd.DataFrame( {col: df[col].mean(axis=1) for col in MODEL_NAMES} ).mul(w_).sum(axis=1)\n\nsubmission.head(N_ROW)\nsubmission.to_csv('submission.csv',index=False)","metadata":{"pycharm":{"name":"#%%\n"},"trusted":true},"execution_count":null,"outputs":[]}]}