{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Context\n\nIn this notebook, I try to test how a good (stability) score can be made by using fixed predictions for each type of discussion.\n\n```\nMAPPING = {\n    'Lead': [0.1, 0.7, 0.2],\n    'Position': [0.1, 0.5, 0.4],\n    'Claim': [0.1, 0.7, 0.2],\n    'Evidence': [0.3, 0.4, 0.3],\n    'Rebuttal': [0.4, 0.2, 0.4],\n    'Counterclaim': [0.1, 0.7 ,0.2],\n    'Concluding Statement': [0.1, 0.6, 0.3]\n}\n```\n\n### How predictions were made\n\n```\nget_preds()\n>>> [0.34, 0.44, 0.22]\n```\n\nTo do this, I randomly generated predictions and tried to find the optimal combination of them.\n\n**TEST | Fast fixed predictions [Not Train]**  \nhttps://www.kaggle.com/code/renokan/test-fast-fixed-predictions-not-train\n\n```\ndef get_preds(n_preds=3, n_round=2):\n    result = []\n    value = 1\n    \n    for _ in range(n_preds-1): \n        if value > 0:\n            x = round(random.uniform(0, value), n_round)\n            value = round(value - x, n_round)\n        else:\n            x = 0\n            \n        result.append(x)\n    \n    result.append(value)\n        \n    return result\n\n[...]\n\nTYPES = train_origin['discourse_type'].unique()\n\nsteps = 450000  # 300000 / 600000\ncheckpoints = [10000, 100000, 150000, 200000, 250000, 300000, 400000, 500000]\n\ntext_col = \"Type\"\ntarget_col = \"Target\"\nx_train = train_data[text_col]\ny_train = train_data[target_col]\n\nbest_scores = {}\n\nfor step in range(1, steps):\n    random_preds = {type_name: get_preds(n_round=1) for type_name in TYPES}\n    \n    predictions = fill_preds(x_train, random_preds)\n    score = get_score(y_train, predictions)\n\n    if step in checkpoints:\n        best_score = sorted(best_scores.keys())\n        checkpoint = f\"checkpoint: {step:<8}\"\n\n    [...]\n```","metadata":{}},{"cell_type":"markdown","source":"# 1. Import & Def & Set","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfrom sklearn.metrics import log_loss","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-13T00:40:09.575648Z","iopub.execute_input":"2022-07-13T00:40:09.576103Z","iopub.status.idle":"2022-07-13T00:40:10.718593Z","shell.execute_reply.started":"2022-07-13T00:40:09.576009Z","shell.execute_reply":"2022-07-13T00:40:10.717626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### What is Log Loss?\n> The likelihood function answers the question \"How likely did the model think the actually observed set of outcomes was\". If that sounds confusing, an example should help.  \nhttps://www.kaggle.com/code/dansbecker/what-is-log-loss","metadata":{}},{"cell_type":"code","source":"def fill_preds(data, mapping):\n    predictions = []\n    for x in data:\n        predictions.append(mapping.get(x))\n        \n    return np.array(predictions)\n\n\ndef get_score(y_true, predictions, n_round=3):\n    result = log_loss(y_true, predictions)\n    \n    return round(result, n_round)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T00:40:10.720400Z","iopub.execute_input":"2022-07-13T00:40:10.721040Z","iopub.status.idle":"2022-07-13T00:40:10.726188Z","shell.execute_reply.started":"2022-07-13T00:40:10.721006Z","shell.execute_reply":"2022-07-13T00:40:10.725494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"LABELS = ['Ineffective', 'Adequate', 'Effective']\nTARGETS = {name: i for i, name in enumerate(LABELS)}\nMAPPING = {\n    'Lead': [0.1, 0.6, 0.3],\n    'Position': [0.1, 0.5, 0.4],\n    'Claim': [0.1, 0.6, 0.3],\n    'Evidence': [0.2, 0.4, 0.4],\n    'Rebuttal': [0.6, 0.3, 0.1],\n    'Counterclaim': [0.1, 0.7, 0.2],\n    'Concluding Statement': [0.3, 0.4, 0.3]\n}\n\nN_ROW = 10","metadata":{"execution":{"iopub.status.busy":"2022-07-13T00:40:10.728001Z","iopub.execute_input":"2022-07-13T00:40:10.728665Z","iopub.status.idle":"2022-07-13T00:40:10.739119Z","shell.execute_reply.started":"2022-07-13T00:40:10.728616Z","shell.execute_reply":"2022-07-13T00:40:10.737855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Check mapping (train data)","metadata":{}},{"cell_type":"code","source":"use_path = \"../input/feedback-prize-effectiveness/train.csv\"\nuse_cols = ['discourse_id', 'discourse_type', 'discourse_effectiveness']\n\ntrain = pd.read_csv(use_path, usecols=use_cols)\n\ntrain.head(N_ROW)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T00:40:10.741640Z","iopub.execute_input":"2022-07-13T00:40:10.742313Z","iopub.status.idle":"2022-07-13T00:40:11.040106Z","shell.execute_reply.started":"2022-07-13T00:40:10.742278Z","shell.execute_reply":"2022-07-13T00:40:11.038852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[LABELS] = pd.DataFrame(\n    fill_preds(train['discourse_type'], MAPPING),\n    columns=LABELS\n)\n\ntrain['Target'] = train['discourse_effectiveness'].replace(TARGETS)\n\ntrain.head(N_ROW)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T00:40:11.041711Z","iopub.execute_input":"2022-07-13T00:40:11.041992Z","iopub.status.idle":"2022-07-13T00:40:11.109783Z","shell.execute_reply.started":"2022-07-13T00:40:11.041966Z","shell.execute_reply":"2022-07-13T00:40:11.108681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_score(train['Target'], train[LABELS])","metadata":{"execution":{"iopub.status.busy":"2022-07-13T00:40:11.111128Z","iopub.execute_input":"2022-07-13T00:40:11.111779Z","iopub.status.idle":"2022-07-13T00:40:11.145082Z","shell.execute_reply.started":"2022-07-13T00:40:11.111737Z","shell.execute_reply":"2022-07-13T00:40:11.143922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Statistics and notes\n\nVersion | Check | Public Score | Split | Note\n-- | -- | -- | -- | --\nV1 | 1.014 | 0.993 | type-label | Score: 1.014 / Valid: 1.014\nV2 | 1.008 | 0.994 | type-label | Score: 1.008 / Valid: 1.008\nV3 | 1.003 | 0.994 | type | Score: 1.003 / Valid: 1.005\nV4 | 1.023 | 1.012 | essay_id | -\nV5 | 1.013 | 0.998 | essay_id | Score: 1.012 / Valid: 1.007\nV6 | 1.019 | 1.004 | essay_id | Score: 1.018 / Valid: 1.018\nV7 | 1.002 | 0.989 | label | Score: 1.001 / Valid: 1.005\nV8 | 1.025 | 0.995 | label | Score: 1.025 / Valid: 1.023\nV9 | 1.025 | 0.996 | type | Score: 1.025 / Valid: 1.028\nV10 | 1.008 | 0.994 | type-label | Score: 1.008 / Valid: 1.008\nV11 | 1.009 | 0.980 | type-label | Score: 1.009 / Valid: 1.009\nV12 | 1.011 | 1.006 | type-label | Score: 1.011 / Valid: 1.011\n\n### How data was split\n\n```\n[...]\nhow_split = 1\n\nif how_split == 1:\n    df['SplitBy'] = df['Type'] + '-' + df['Label']\n    \nelif how_split == 2:\n    df['SplitBy'] = df['Type']\n    \nelif how_split == 3:\n    df['SplitBy'] = train_origin['essay_id']\n\nelse:\n    df['SplitBy'] = df['Label']\n[...]\n\ntrain_indx, val_indx = train_test_split(\n    df.index, stratify=df['SplitBy'].values,\n    test_size=0.15, random_state=2022\n)\n\ncols_list = ['Type', 'Label', 'Target']\ntrain_data = df.loc[train_indx, cols_list].copy().reset_index(drop=True)\nvalid_data = df.loc[val_indx, cols_list].copy().reset_index(drop=True)\n```","metadata":{}},{"cell_type":"markdown","source":"# 3. Create submission","metadata":{}},{"cell_type":"code","source":"use_path = \"../input/feedback-prize-effectiveness/sample_submission.csv\"\n\nsubmission = pd.read_csv(use_path)\n\nsubmission.head(N_ROW)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T23:39:18.564844Z","iopub.execute_input":"2022-07-11T23:39:18.565937Z","iopub.status.idle":"2022-07-11T23:39:18.587422Z","shell.execute_reply.started":"2022-07-11T23:39:18.565889Z","shell.execute_reply":"2022-07-11T23:39:18.586049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"use_path = \"../input/feedback-prize-effectiveness/test.csv\"\nuse_cols = ['discourse_id', 'discourse_type']\n\ntest = pd.read_csv(use_path, usecols=use_cols)\n\ntest.head(N_ROW)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T23:39:18.616587Z","iopub.execute_input":"2022-07-11T23:39:18.616976Z","iopub.status.idle":"2022-07-11T23:39:18.635344Z","shell.execute_reply.started":"2022-07-11T23:39:18.616945Z","shell.execute_reply":"2022-07-11T23:39:18.634095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission[LABELS] = pd.DataFrame(\n    fill_preds(test['discourse_type'], MAPPING),\n    columns=LABELS\n)\n\nsubmission.head(N_ROW)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T23:39:18.661506Z","iopub.execute_input":"2022-07-11T23:39:18.661899Z","iopub.status.idle":"2022-07-11T23:39:18.679614Z","shell.execute_reply.started":"2022-07-11T23:39:18.661868Z","shell.execute_reply":"2022-07-11T23:39:18.678137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T23:39:18.701597Z","iopub.execute_input":"2022-07-11T23:39:18.701993Z","iopub.status.idle":"2022-07-11T23:39:18.712265Z","shell.execute_reply.started":"2022-07-11T23:39:18.701960Z","shell.execute_reply":"2022-07-11T23:39:18.711300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}