{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\">Cost Minimization & Floor - Part(1)</h1></h1>\n    <h2 align=\"center\">Identify the position of a smartphone in a shopping mall</h2>\n    <h3 align=\"center\">By: Somayyeh Gholami & Mehran Kazeminia</h3>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Description:","metadata":{}},{"cell_type":"markdown","source":"### - In this notebook (No. 1), we used the following magic notebook for \"Cost Minimization\".\n\nhttps://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\n\n### - We used the following creative notebook for \"Fix the floor prediction\".\n\nhttps://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\n\n### - In this notebook (No. 1), we improve the results of ten public notebooks by the above methods. Then in the notebook (No. 2) we will use \"Ensembling\" and \"Comparative Method\". Finally, the so-called \"Snap to Grid\" notebook (No. 3) produces the final result. Thanks to everyone who shared their notebooks, the addresses of some of the used notebooks are as follows:\n\nhttps://www.kaggle.com/robikscube/indoor-navigation-snap-to-grid-post-processing\n\nhttps://www.kaggle.com/therocket290/lstm-unified-wi-fi-training-x-and-y-with-floor\n\nhttps://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats\n\nhttps://www.kaggle.com/oxzplvifi/indoor-gbm-postprocessing-xy-prediction\n\nhttps://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\n\nhttps://www.kaggle.com/hiro5299834/wifi-features-with-lightgbm-kfold\n\nhttps://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction\n\n\n### - As we explained in previous notebooks; If you upgrade the score of all notebooks with the \"Snap to Grid\" method before \"Ensembling\" and then perform the \"Ensembling\" operation, all the errors will add up and you will not get a good result. This means using the \"Snap to Grid\" method only in the last step. But \"Cost Minimization\" can be done before or after \"Ensembling\". Of course, as you can see, we do \"Cost Minimization\" for all results from the beginning.\n\n### =======================================================\n\n### For more information, you can refer to the following address:\n\nhttps://www.kaggle.com/c/indoor-location-navigation/discussion/230153\n\n## >>> Good Luck <<<\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# If you find this work useful, please don't forget upvoting :)","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Import & Data Set","metadata":{}},{"cell_type":"code","source":"!git clone --depth 1 https://github.com/location-competition/indoor-location-competition-20 indoor_location_competition_20\n    \n!rm -rf indoor_location_competition_20/data","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport scipy.sparse\nimport scipy.interpolate\n\nfrom tqdm import tqdm\nimport multiprocessing\nimport matplotlib.pyplot as plt\n\nfrom indoor_location_competition_20.io_f import read_data_file\nimport indoor_location_competition_20.compute_f as compute_f\n\n%matplotlib inline","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"INPUT_PATH = '../input/indoor-location-navigation'","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Kernels Data (Public Score & File Path)\n\ndfk = pd.DataFrame({ \n    'Kernel ID': ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J'],  \n    'Score':     [9.773, 9.530, 8.500, 8.418, 8.333, 8.073, 7.745, 7.661, 7.274, 7.285],   \n    'File Path': ['../input/indoor9773/Indoor9773.csv', '../input/indoor9530/Indoor9530.csv', '../input/indoor8500/Indoor8500.csv', '../input/indoor8418/Indoor8418.csv', '../input/indoor8333/Indoor8333.csv', '../input/indoor8073/Indoor8073.csv', '../input/indoornav7745sub/submission.csv', '../input/indoor7661/Indoor7661.csv', '../input/indoor-wifi-floor/submission.csv', '../input/time-series-rnn-xy-prediction/submission.csv']     \n})    \n    \ndfk","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Functions\n\nThese descriptions and codes are copied from the following notebook:\n\nhttps://www.kaggle.com/saitodevel01/indoor-post-processing-by-cost-minimization\n","metadata":{}},{"cell_type":"markdown","source":"To combine machine learning (wifi features) predictions with sensor data (acceleration, attitude heading),\nI defined cost function as follows,\n$$\nL(X_{1:N}) = \\sum_{i=1}^{N} \\alpha_i \\| X_i - \\hat{X}_i \\|^2 + \\sum_{i=1}^{N-1} \\beta_i \\| (X_{i+1} - X_{i}) - \\Delta \\hat{X}_i \\|^2\n$$\nwhere $\\hat{X}_i$ is absolute position predicted by machine learning and $\\Delta \\hat{X}_i$ is relative position predicted by sensor data.\n\nSince the cost function is quadratic, the optimal $X$ is solved by linear equation $Q X = c$\n, where $Q$ and $c$ are derived from above cost function.\nBecause the matrix $Q$ is tridiagonal,\neach machine learning prediction is corrected by *all* machine learning predictions and sensor data.\n\nThe optimal hyperparameters ($\\alpha$ and $\\beta$) can be estimated by expected error of machine learning and sensor data,\nor just tuned by public score.","metadata":{}},{"cell_type":"code","source":"def compute_rel_positions(acce_datas, ahrs_datas):\n    \n    step_timestamps, step_indexs, step_acce_max_mins = compute_f.compute_steps(acce_datas)\n    headings = compute_f.compute_headings(ahrs_datas)\n    stride_lengths = compute_f.compute_stride_length(step_acce_max_mins)\n    step_headings = compute_f.compute_step_heading(step_timestamps, headings)\n    rel_positions = compute_f.compute_rel_positions(stride_lengths, step_headings)\n    \n    return rel_positions","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def correct_path(args):\n    path, path_df = args\n    \n    T_ref  = path_df['timestamp'].values\n    xy_hat = path_df[['x', 'y']].values\n    \n    example = read_data_file(f'{INPUT_PATH}/test/{path}.txt')\n    rel_positions = compute_rel_positions(example.acce, example.ahrs)\n    if T_ref[-1] > rel_positions[-1, 0]:\n        rel_positions = [np.array([[0, 0, 0]]), rel_positions, np.array([[T_ref[-1], 0, 0]])]\n    else:\n        rel_positions = [np.array([[0, 0, 0]]), rel_positions]\n    rel_positions = np.concatenate(rel_positions)\n    \n    T_rel = rel_positions[:, 0]\n    delta_xy_hat = np.diff(scipy.interpolate.interp1d(T_rel, np.cumsum(rel_positions[:, 1:3], axis=0), axis=0)(T_ref), axis=0)\n\n    N = xy_hat.shape[0]\n    delta_t = np.diff(T_ref)\n    alpha = (8.1)**(-2) * np.ones(N)\n    beta  = (0.30 + 0.30 * 1e-3 * delta_t)**(-2)\n    A = scipy.sparse.spdiags(alpha, [0], N, N)\n    B = scipy.sparse.spdiags( beta, [0], N-1, N-1)\n    D = scipy.sparse.spdiags(np.stack([-np.ones(N), np.ones(N)]), [0, 1], N-1, N)\n\n    Q = A + (D.T @ B @ D)\n    c = (A @ xy_hat) + (D.T @ (B @ delta_xy_hat))\n    xy_star = scipy.sparse.linalg.spsolve(Q, c)\n\n    return pd.DataFrame({\n        'site_path_timestamp' : path_df['site_path_timestamp'],\n        'floor' : path_df['floor'],\n        'x' : xy_star[:, 0],\n        'y' : xy_star[:, 1],\n    })\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: A\n\nhttps://www.kaggle.com/deepijongwonkim/wifi-features-neural-networks-starter\n\nPublic Score: 9.773","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[0, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('a9773cm99.csv', index=False)\n\na9773cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimization\" & \"Fix the floor prediction\"\n\n### a9773cm99.csv |  Public Score: 7.110","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: B\n\nhttps://www.kaggle.com/byfone/indoor-location-wi-fi-features-catboost-starter\n\nPublic Score: 9.530","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[1, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('b9530cm99.csv', index=False)\n\nb9530cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimization\" & \"Fix the floor prediction\"\n\n### b9530cm99.csv |  Public Score: 6.674","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: C\n\nhttps://www.kaggle.com/hiro5299834/wifi-features-with-lightgbm-groupkfold\n\nPublic Score: 8.500","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[2, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('c8500cm99.csv', index=False)\n\nc8500cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimization\" & \"Fix the floor prediction\"\n\n### c8500cm99.csv |  Public Score: 6.290\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: D\n\nhttps://www.kaggle.com/hiro5299834/wifi-features-with-lightgbm-and-xgboost-kfold\n\nPublic Score: 8.418","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[3, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('d8418cm99.csv', index=False)\n\nd8418cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimization\" & \"Fix the floor prediction\"\n\n### d8418cm99.csv |  Public Score: 6.189","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: E\n\nhttps://www.kaggle.com/hiro5299834/wifi-features-with-lightgbm-kfold\n\nPublic Score: 8.333","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[4, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('e8333cm99.csv', index=False)\n\ne8333cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimization\" & \"Fix the floor prediction\"\n\n### e8333cm99.csv |  Public Score: 6.077\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: F\n\nhttps://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model\n\nPublic Score: 8.073","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[5, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"# simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\n# sub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('f8073cm99.csv', index=False)\n\nf8073cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimization\" & \"Fix the floor prediction\"\n\n### f8073cm99.csv |  Public Score: 6.062","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: G\n\nhttps://www.kaggle.com/oxzplvifi/indoor-gbm-postprocessing-xy-prediction\n\nPublic Score: 7.745","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[6, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('g7745cm99.csv', index=False)\n\ng7745cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimizat\" & \"Fix the floor prediction\"\n\n### g7745cm99.csv |  Public Score: 5.995","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: H\n\nhttps://www.kaggle.com/kokitanisaka/lstm-by-keras-with-unified-wi-fi-feats\n\nPublic Score: 7.661","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[7, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('h7661cm99.csv', index=False)\n\nh7661cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimizat\" & \"Fix the floor prediction\"\n\n### h7661cm99.csv |  Public Score: 5.694\n","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: I\n\nhttps://www.kaggle.com/therocket290/lstm-unified-wi-fi-training-x-and-y-with-floor\n\nPublic Score: 7.274","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[8, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('i7274cm99.csv', index=False)\n\ni7274cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimizat\" & \"Fix the floor prediction\"\n\n### i7274cm99.csv |  Public Score: 5.471","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Kernel: J\n\nhttps://www.kaggle.com/ebinan92/time-series-rnn-xy-prediction\n\nPublic Score: 7.285","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(dfk.iloc[9, 2])\n\ntmp = sub['site_path_timestamp'].apply(lambda s : pd.Series(s.split('_')))\nsub['site'] = tmp[0]\nsub['path'] = tmp[1]\nsub['timestamp'] = tmp[2].astype(float)\n\nprocesses = multiprocessing.cpu_count()\nwith multiprocessing.Pool(processes=processes) as pool:\n    dfs = pool.imap_unordered(correct_path, sub.groupby('path'))\n    dfs = tqdm(dfs)\n    dfs = list(dfs)    \nsub = pd.concat(dfs).sort_values('site_path_timestamp')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fix the floor prediction","metadata":{}},{"cell_type":"code","source":"simple_accurate_99 = pd.read_csv(dfk.iloc[5, 2])\n\nsub['floor'] = simple_accurate_99['floor'].values\n\nsub.to_csv('j7285cm99.csv', index=False)\n\nj7285cm99 = sub","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### After \"Cost Minimizat\" & \"Fix the floor prediction\"\n\n### j7285cm99.csv |  Public Score: 5.847","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Results","metadata":{}},{"cell_type":"code","source":"gfk = pd.DataFrame({ \n    \n    'Kernel ID'   : ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J'],  \n    \n    'Before Score': [9.773, 9.530, 8.500, 8.418, 8.333, 8.073, 7.745, 7.661, 7.274, 7.285], \n    \n    'After Score' : [7.110, 6.674, 6.290, 6.189, 6.077, 6.062, 5.995, 5.694, 5.471, 5.847], \n    \n    'File Name'   : ['a9773cm99.csv', 'b9530cm99.csv', 'c8500cm99.csv', 'd8418cm99.csv', 'e8333cm99.csv', 'f8073cm99.csv', 'g7745cm99.csv', 'h7661cm99.csv', 'i7274cm99.csv', 'j7285cm99.csv'] \n    \n})    \n    \ngfk","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"sub = i7274cm99\nsub.to_csv(\"submission.csv\", index=False)\n\na9773cm99.to_csv(\"a9773cm99.csv\", index=False)\nb9530cm99.to_csv(\"b9530cm99.csv\", index=False)\nc8500cm99.to_csv(\"c8500cm99.csv\", index=False)\nd8418cm99.to_csv(\"d8418cm99.csv\", index=False)\ne8333cm99.to_csv(\"e8333cm99.csv\", index=False)\nf8073cm99.to_csv(\"f8073cm99.csv\", index=False)\ng7745cm99.to_csv(\"g7745cm99.csv\", index=False)\nh7661cm99.to_csv(\"h7661cm99.csv\", index=False)\ni7274cm99.to_csv(\"i7274cm99.csv\", index=False)\nj7285cm99.to_csv(\"j7285cm99.csv\", index=False)\n\n!ls","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}}]}