{"cells":[{"metadata":{"_kg_hide-input":false},"cell_type":"markdown","source":"In order to switch gears to a classification task, I decided to explore how to discretize time_to_failure values.\nIn particular, I was wonder how many different deltas, and how they distributed through the time.\n\nI also will use the terms from the reinforcement learning paradigm, like a reward or an episode."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"import os\nimport csv\nimport numpy as np\nfrom tqdm import tqdm\nfrom collections import Counter\nfrom operator import itemgetter\nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Read 629145480 time_to_failure values from train.csv"},{"metadata":{"trusted":true},"cell_type":"code","source":"N = 629145480\ny = np.empty(N, dtype=np.float32)\n\nwith open('../input/train.csv') as f:\n    reader = csv.reader(f)\n    for header in reader:\n        break\n    for i, row in enumerate(tqdm(reader, total=N)):\n        y[i] = float(row[1])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Determine number of episodes"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"deltas = y[1:] - y[:-1]\nepisodes = {0, len(y)}\nepisodes.update(np.arange(0, len(y) - 1)[deltas > 0] + 1)\nepisodes = sorted(list(episodes))\nepisodes = list(zip(episodes[:-1], episodes[1:]))\nprint('Episodes:', len(episodes))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Determine unique reward values by constraining all values with round() function"},{"metadata":{"_kg_hide-output":true,"_kg_hide-input":false,"trusted":true},"cell_type":"code","source":"deltas = []\nrewards = [] # non-zero deltas\n\nfor start, end in episodes:\n    t = y[start:end]\n    d = t[1:] - t[:-1]\n    d = np.round(d, 10)\n    deltas.append(d)\n    rewards.extend(d[d != 0])\n\ncounts = Counter(rewards)\n\nclasses = dict()\nfor value, numbers in sorted(counts.items(), key=itemgetter(0)):\n    print('%2d % .10f %7d' % (len(classes), value, numbers))\n    classes[value] = len(classes)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"At this point we have 52 different reward values and how many times they occur in the training data. But it doesn't tell how often the values change. Thus, I decided to fill the intermediate regions with a following non-zero reward label."},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"targets = [] # labels will store in a reverse order\nfor i, d in enumerate(deltas[::-1]):\n    c = 0\n    for r in tqdm(d[::-1], desc='Episode %d' % i):\n        if r != 0:\n            c = classes[r] if r in classes else 0\n        targets.append(c)\n    targets.append(c)\n# reverse again to match an original order\ntargets = np.array(targets[::-1], dtype=np.int8)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Targets values are highly unbalanced"},{"metadata":{"trusted":true},"cell_type":"code","source":"counts = np.bincount(targets, minlength=52)\nplt.bar(range(52), counts)\nplt.xlabel(\"label\")\nplt.ylabel(\"number of samples\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Eliminate non-valuable classes with less than 10M samples"},{"metadata":{"trusted":true},"cell_type":"code","source":"n_rewards = 1\nmapping = {-1: 0}\n\nfor i, c in enumerate(counts):\n    if c < 10_000_000:\n        mapping[i] = 0\n        continue\n    print(n_rewards, i, c)\n    mapping[i] = n_rewards\n    n_rewards += 1\n\ntargets_balanced = targets.copy()\nfor k, v in mapping.items():\n    targets_balanced[targets == k] = v","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Finally, targets_balanced array could be used for a classification task with 9 classes. Unfortunatly, I could not build a usefull model. On the other hand, visualization of targets values gives some insight into the training data."},{"metadata":{"trusted":true},"cell_type":"code","source":"x = np.arange(0, len(y))\n\n_x = x[::1000]\n_y = y[::1000]\n_c = targets_balanced[::1000]\n\nsc = plt.scatter(_x, _y, c=_c, s=50, cmap='jet')\nplt.colorbar(sc)\nplt.xlabel(\"time\")\nplt.ylabel(\"time_to_failure\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Final remarks:\n* all my code of filtering unique target values seems useless, because it could be easily determined as a power of 2\n* from my point of view, it looks very artificial and discrete, perhaps it explains the underlain process behind the lab equipment\n\nSuggestions for future work:\n* make a stratification strategy based on different color regions\n* build a classifier and used it in a balancing procedure for the test data"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}