{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Consideration of floor\n## Summary\nA BBSID is an ID attached to a physical equipment of WiFi access point. Since WiFI access points are fixed to the building, they do not\nmove. Therefore a BSSID, in whichever path it appears in the dataset, is uniquely tied to one floor number.\n\nWe know from the training data, who (= a person on the path) received a WiFi signal from a particular WiFi access point (= BSSID. hereafter we use 'BBSID' and 'WiFi access point' interchangeably). We also know from the training data, on which floor this person was at the time of the signal reception. In this way, we can assign one floor to a particular BSSID.\n\nOne problem is that a person sometimes pick up a WiFi signal from the floors on which the person is not located, because WiFi signals can go through walls and floors, and a part of the building where the floor is removed to connect the adjacent stories.\n\nHowever, if we select those signal-reception events with strong signals (= large RSSI), the person should be close to the WiFi access point, and therefore the chances are high that the person is on the same floor with this WiFi access point. By selectively looking at the floor information in such large-signal reception-events, we can tie a BSSID to a particular floor with better accuracy.\n\nOnce we created a table of BSSIDs with the information on which floor they are in the building, we can use the table to assign a floor to a path in the test data by looking at the set of BSSIDs that are show up along the path.\n\nAgain the set of BSSIDs in a particular path could be contaminated by WiFi access points on other floors. We again selectively look at the signal-reception events with strong signals (= large RSSI) to filter out unexpected signal receptions from other floors.\n\nIn this notebook, the data analysis goes as follows.\n\n1. We pick up one site. \n2. List up all BSSID-RSSI pairs (=reception events), regardless on which paths they happen.\n3. Group the list by BSSID, and select 8 pairs (the number is arbitrary) that have strongest RSSI.\n4. Look at on which floor the pair happens.\n5. Now one BSSID has 8 entires of the floor. Let them vote, and find the floor that occurred most often. Assign this floor as the location of this BSSID.\n6. Use a validation set, and predict the floor of a particular path, using their BSSIDs that show up along the path.\n\n## Conclusion\nThere is one parameter we had to adjust, which is the number of signal-reception events to be included in the procedure 6. Once the number is tweaked, the prediction of the floor was perfect for the first building we tried (= zero error).\n \nOptimal number of votes that should be considered in the procedure 6 likely differs building by building. It makes sense as the range of WiFi signal differs depending on the structure of a building. It is worthwhile tuning the parameter, because the errors in the floor predictiosn cost more than those in the two dimensional positions.\n\n## Acknowledgement\nI used the data published by @kokitanisaka\nhttps://www.kaggle.com/kokitanisaka/create-unified-wifi-features-example\n\nThis notebook is inspired by the work of @nigelhenly \nhttps://www.kaggle.com/nigelhenry/simple-99-accurate-floor-model"},{"metadata":{},"cell_type":"markdown","source":"---\n## Data Processing"},{"metadata":{},"cell_type":"markdown","source":"Set up packages, notebook-wide parameters and data directory."},{"metadata":{"trusted":true,"collapsed":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom pathlib import Path\nimport os\nfrom datetime import datetime\n\nfrom sklearn.model_selection import StratifiedKFold, GroupKFold\nfrom sklearn.preprocessing import StandardScaler, LabelEncoder\n\nfrom datatable import (dt, f, by)\nfrom IPython.core.display import display, HTML\n\ndisplay(HTML(\n    '<style>'\n        '#notebook { padding-top:0px !important; } '\n        '.container { width:100% !important; } '\n        '.end_space { min-height:0px !important; } '\n        '</style>'\n        ))\n\n# ========================================================\npd.options.display.max_rows = 999\ndt.options.display.max_nrows = 999\n# ========================================================\n\npath = Path('../input/indoorunifiedwifids')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Read the data creatd by @kokitanisaka. Let us use datatable, instead of pandas, for the speed. "},{"metadata":{"trusted":true},"cell_type":"code","source":"# ========================================================\n# reading data\n\ndata = dt.fread(path/'train_all.csv')\ntest_data = dt.fread(path/'test_all.csv')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The content of `data` looks like this."},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"code","source":"data","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"It is too cumbersome to work with long text. The strings are converted to numbers. First, we need the name of the columns in the datatable."},{"metadata":{"trusted":true},"cell_type":"code","source":"bx = [i for i in data.names if i.startswith('bssid_')]\nrx = [i for i in data.names if i.startswith('rssi_')]\nbx, rx\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Collect all BSSID names that occur in the training and the test dataset."},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"code","source":"wifi_bssids = dt.unique(dt.rbind(data[:, bx], test_data[:, bx]))\nwifi_bssids","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"wifi_bssids.nrows","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"There are 61307 unique BSSIDs in the dataset. Prepare label encoders that convert strings to numbers. Apply the encoders on `data`."},{"metadata":{"trusted":true},"cell_type":"code","source":"le = LabelEncoder()\nle.fit(wifi_bssids)\n\nle_site = LabelEncoder()\nle_site.fit(data[:, 'site_id'])\n\nle_path = LabelEncoder()\nle_path.fit(data[:, 'path'])\n\n# ========================================================\n\ndata[:, 'site_id'] = le_site.transform(data[:, 'site_id'])\ndata[:, 'path'] = le_path.transform(data[:, 'path'])\n\nfor i in bx:\n    data[:, i] = le.transform(data[:, i])","execution_count":null,"outputs":[]},{"metadata":{"scrolled":false},"cell_type":"markdown","source":"Now the data looks easer to read."},{"metadata":{"trusted":true},"cell_type":"code","source":"# from sklearn.model_selection import StratifiedKFold, GroupKFold\ndata\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.names","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"__Let us pick up one site.__ When we put the notebook in production, here is the part we have to change, and loop. We pick up the first site, where `data['site'] == 0`"},{"metadata":{"trusted":true},"cell_type":"code","source":"d0 = data[f.site_id==0, :]\nd0","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"BSSID part of data."},{"metadata":{"trusted":true},"cell_type":"code","source":"d0[:,bx]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"RSSI part of data."},{"metadata":{"trusted":true},"cell_type":"code","source":"d0[:,rx]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now we make a pair of these two tables. It is like overlapping one table on the other, and make a pair. The resultant list will be long. The number of columns (`rx` and `bx`) are 100, and the number of rows in data is 9296. The expected length of the list is therefore 929600."},{"metadata":{"trusted":true},"cell_type":"code","source":"len(bx), len(rx), d0.nrows","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"tp1 = datetime.now()\ntpx0 = datetime.now()\n\nfor i in range(d0.nrows):\n    dbx = d0[i,bx]\n    drx = d0[i,rx]\n\n    dbx = dt.Frame(dbx.to_numpy().T)\n    drx = dt.Frame(drx.to_numpy().T)\n#    print(f'dbx {dbx.nrows}') \n\n    dfx = dt.repeat(dt.Frame({'floor':[d0[i,'floor']]}), dbx.nrows)\n    dpx = dt.repeat(dt.Frame({'path':[d0[i,'path']]}), dbx.nrows)\n\n    cx = dt.cbind([dbx, drx, dfx, dpx])\n    if i ==0 :  \n        dx = cx\n    else:\n#        dx = dt.rbind([cx, dx])\n        dx.rbind(cx)\n        \n    if i % 1000 == 0:\n        tpx = datetime.now()\n        print(f'{i:-5} \\033[32m{tpx}\\033[0m {dx.nrows:-8} \\033[91m{tpx-tpx0}\\033[0m')\n        tpx0 = tpx\n\ndx.names = ['BSSID', 'RSSI', 'floor', 'path']    \ntp2 = datetime.now()\n\nprint(f'took \\033[1,32m{tp2-tp1}\\033[0m')\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dx","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We convert dx to a pandas data frame now here."},{"metadata":{"trusted":true},"cell_type":"code","source":"df = dx.sort([f.BSSID, f.RSSI]).to_pandas()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The data frame will be split to training (`ti`) and validation sets (`vi`). "},{"metadata":{"trusted":true},"cell_type":"code","source":"N_SPLIT=5\ngkf = GroupKFold(n_splits=N_SPLIT)\nti, vi = next(gkf.split(df, df['floor'], groups=df['path']))\ndf_ti = df.iloc[ti]\ndf_vi = df.iloc[vi]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Make a list of unique BSSIDs."},{"metadata":{"trusted":true},"cell_type":"code","source":"u_bssid = df_ti['BSSID'].unique()\nu_bssid, len(u_bssid)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us look at the first 3 BSSIDS, 8, 107, and 129. "},{"metadata":{"trusted":true},"cell_type":"code","source":"x_df = pd.DataFrame(columns=['BSSID_S', 'BSSID', 'floor'])\nfor i in u_bssid[:3]:\n\n    print(f'\\033[31mBSSID \\033[0m {i:5}')\n\n    x = df_ti.loc[df_ti['BSSID']==i].sort_values(['RSSI'], ascending=False).iloc[0:8]\n    xb = le.inverse_transform([i])[0]\n\n    x_floor = x['floor'].mode().values[0]\n    x_a = pd.DataFrame({'BSSID_S':[xb], 'BSSID':[i], 'floor':[x_floor]})\n\n    print(x)\n    print(f'\\033[32mfloor voted: \\033[0m{x_floor}')\n    print()\n\n    if i == u_bssid[0]: \n        x_df = x_a\n    else: \n        x_df = pd.concat([x_df, x_a])\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In case of the BSSID number 8, there are 4 different paths in data (10844, 10728, 10734, 10733) among 8 events that received strong signals from this WiFi access point. All voted for floor 0. We conclude BSSID #8 is locate on the first floor, F1. The voting for the WiFi access point BSSID #129 is not unanimous, but according to the 8 strongest RSSIs receptions, it is 1. Note that this number '8 strongest signal receptions' is an arbitrary choice, and there is a room of an experiment.  "},{"metadata":{},"cell_type":"markdown","source":"We will run the code for whole `u_bssid`. The text BSSID names ('BSSID_S') are added. "},{"metadata":{"trusted":true},"cell_type":"code","source":"x_df = pd.DataFrame(columns=['BSSID_S', 'BSSID', 'floor'])\ntp1 = datetime.now()\ntpx0 = datetime.now()\n\nfor i_b, i in enumerate(u_bssid):\n\n    xb = le.inverse_transform([i])[0]\n    x0 = df_ti.loc[(df_ti['BSSID'] == i) & (df_ti['RSSI'] != -999)]\n\n    if len(x0) == 0:\n#        x_floor = np.nan\n         x_floor = 999\n\n    else: \n        p = np.min([8, len(x0)])\n\n        x  = x0.sort_values(['RSSI'], ascending=False).iloc[0:p]\n        x_floor = x['floor'].mode().values[0]\n\n    x_a = pd.DataFrame({'BSSID_S':[xb], 'BSSID':[i], 'floor':[x_floor]})\n\n    if i_b == 0: \n        x_df = x_a\n    else: \n        x_df = pd.concat([x_df, x_a])\n\n    if i_b % 1000 == 0:\n        tpx = datetime.now()\n        print(f'{i:-8} \\033[32m{tpx}\\033[0m {len(x_df): 8} \\033[91m{tpx-tpx0}\\033[0m')\n        tpx0 = tpx\n\ntp2 = datetime.now()\nprint(f'took \\033[1,32m{tp2-tp1}\\033[0m')\n\n\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The contents of `x_df`."},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"code","source":"x_df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(x_df)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This is the table that locates each BSSID on a certain floor. "},{"metadata":{},"cell_type":"markdown","source":"Now we will start working with the validation set. First, collect all unique paths in the validation dataset. "},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"code","source":"u_path = df_vi['path'].unique()\nu_path, len(u_path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Just to refresh our memory what was in `df_vi`."},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"code","source":"df_vi","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let us take a look at first 3 paths, 10678, 10659 and 352. "},{"metadata":{"trusted":true},"cell_type":"code","source":"f_p = np.zeros(0)\nf_t = np.zeros(0)\n\nfor i in u_path[:3]: \n    x = df_vi.loc[df['path']==i]\n\n    x_mer = x.merge(x_df.drop(['BSSID_S'], axis=1), left_on='BSSID', right_on='BSSID', suffixes=('','_pred'))\n    x_mer_top = x_mer.sort_values(['RSSI'], ascending=False)[0:p]\n\n    p1 = x_mer_top['floor_pred'].mode().values[0]\n    f_p = np.hstack((f_p, p1))\n    f_t = np.hstack((f_t, x_mer['floor'].mode().values[0]))\n\n    print(f'\\033[31mpath \\033[0m{i:6}\\033[0m')\n    print(f'{x_mer_top}')\n    print(f'\\033[32mfloor voted: \\033[0m{p1}')\n    print()\n","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The BSSIDs observed along the path 10678 all vote for the WiFi device being located on the floor -1, if we picked the 8 strongest signal-receptions. One can safely attribute the path to the floor B1. "},{"metadata":{},"cell_type":"markdown","source":"We will look into how many signal-reception events is the best to come up to the correct assignment of the floor to a path. `p` is changed from 1 to 32 below. With `p=26`, the code makes no mistake in the assignment of the floor."},{"metadata":{"scrolled":false,"trusted":true},"cell_type":"code","source":"for p in np.arange(1,32,1):\n\n    f_p = np.zeros(0)\n    f_t = np.zeros(0)\n\n    for i in  u_path: \n        x = df_vi.loc[df['path']==i]\n\n        x_mer = x.merge(x_df.drop(['BSSID_S'], axis=1), left_on='BSSID', right_on='BSSID', suffixes=('','_pred'))\n        x_mer_top = x_mer.sort_values(['RSSI'], ascending=False)[0:p]\n\n        f_p = np.hstack((f_p, x_mer_top['floor_pred'].mode().values[0]))\n        f_t = np.hstack((f_t, x_mer['floor'].mode().values[0]))\n\n    print(f'{p:-2} {(f_p==f_t).mean():.4f}')\n\n","execution_count":null,"outputs":[]},{"metadata":{"scrolled":true,"trusted":true},"cell_type":"markdown","source":"Note this number `p` varries accoring to which site we are looking at. Sometimes we need `p > 100' to get right answers. \n"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}