{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# A Notebook For Your Augury\nBy: Elliott de Launay\n\nTemplate to make machine learning a bit easier to follow and repeatable.\n\nGithub: [A Notebook For Your Augury](https://github.com/edelauna/a-notebook-for-your-augury)\n\nPython's a funny language, will try out a build as we go kind of style.\n\nHow To Use:\n* Best to go cell by cell\n* Cell's are organized so that at the top there is some kind of `ALL_CAPS` variable which is either an array of Dicts or Dict, with the intent of passing this as input to the functions specified.\n* After the config, code is marked betwen a block:\n```\n  # Some Config here\n  \"\"\"\"\n  \"\"\"\"\n  # There will be some code here which is the mechanics of what's being performed\n  # This code then dynamically added to the Augury class \n  #    (In hopes it makes it easier to follow along)\n  \"\"\"\"\n  \"\"\"\"\n  # Then the actual code which runs the above function with the configs\n```\n\nGoal:\n* Automate:\n  * Transformations applied to Training Data, also get applied test.\n  * Evaluate 3 Models\n    * Included a 4th Model (KMeans) as _experiemental_\n  * Select best performing model to be used for Test","metadata":{"id":"JSeZQwtuSUeu"}},{"cell_type":"code","source":"### Initialization\nclass Augury:\n  def constructor():\n    \"\"\"\n    'Decorator' which adds funcs to this class    \n    \"\"\"\n    def decorator(*funcs):\n      for func in funcs:\n        setattr(__class__, func.__name__, func)\n    return decorator\n  def _attr(self, *readers, overwrite=True, **writers):\n    \"\"\"\n    Returns the relevent self.attr.\n    Will also set a value if passed with a value.\n      Parameters\n      ----------\n      readers : str\n          attr to return\n      overwrite : bool\n          if False, will not set a value, if a value exists\n      writers : Dict\n          key as attr to return, value as value to set.\n      Returns\n      ----------\n      tuple of ordered attrs\n    \"\"\"\n    _attrs = []\n    for attr in readers:\n      _attrs.append(getattr(self,attr)) if hasattr(self, attr) else _attrs.append(None)\n    for attr in writers:\n      if not hasattr(self, attr) or overwrite:\n        setattr(self, attr, writers[attr])\n      _attrs.append(getattr(self,attr))\n    return tuple(_attrs)\n  def _hash(self, hash, *keys):\n    \"\"\"\n    Helper function to destructure a hash.\n    ...didn't end up using this as much\n      Parameters\n      ----------\n      hash : Dictionary\n          hash to destructure\n      keys : str\n          values at key will be returned.\n      Returns\n      ----------\n      tuple of ordered hash values\n    \"\"\"\n    results = []\n    for key in keys:\n      results.append(hash[key]) if key in hash else results.append(None)\n    return tuple(results)\n\n# Setting up cache for some longer running functions\nimport os\nAUGURY_CACHE=\".augury/.cache\"\nos.environ[\"AUGURY_CACHE\"]=AUGURY_CACHE\n! mkdir -p $AUGURY_CACHE\n\n# Setting main variable to hold it all\naugur = Augury()","metadata":{"id":"RaQk_59pRwdb","execution":{"iopub.status.busy":"2022-07-09T17:31:18.422373Z","iopub.execute_input":"2022-07-09T17:31:18.422791Z","iopub.status.idle":"2022-07-09T17:31:19.237721Z","shell.execute_reply.started":"2022-07-09T17:31:18.422758Z","shell.execute_reply":"2022-07-09T17:31:19.236534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Pipeline\nThe goal of this notebook is to provide a template for machine learning modelling when working with \"simple\" Continous and Categorical data. The notebook will perform the following:\n1. Read in Data\n2. Explore Data\n    1. Set Label\n    1. Categorical Features\n        1. Handle Nulls\n        2. Feature Engineering\n    3. Continous Features\n        1. Handle Nulls\n        1. Cap and Floor Outliers\n        2. Feature Engineering\n1. Transform\n  1. Transform Skewed Features\n  2. Transform to Numerical Indicators\n  1. Feature Scaling & Optimal Count\n2. Principal Component Analysis\n1. Train / Validation Split\n3. Model Selection\n    1. Hyperparameter Tuning\n    1. Fit\n    1. Evaluation\n4. Test\n","metadata":{"id":"RaY2rbB7UC2E"}},{"cell_type":"markdown","source":"## Read in Data","metadata":{"id":"8S6PONWTju85"}},{"cell_type":"code","source":"\"\"\"\nBeginnin our Analysis\n\"\"\"\nimport pandas as pd\n\n# Confirm this path\nTRAIN_PATH=\"/kaggle/input/spaceship-titanic/train.csv\"\nTEST_PATH=\"/kaggle/input/spaceship-titanic/test.csv\"\n\ntraining_data = pd.read_csv(TRAIN_PATH)\ntraining_data.head()","metadata":{"id":"usWs1OW2sbFc","outputId":"2f1a0626-70da-4687-89ee-8d45cdaeccfa","execution":{"iopub.status.busy":"2022-07-09T17:31:19.241243Z","iopub.execute_input":"2022-07-09T17:31:19.241726Z","iopub.status.idle":"2022-07-09T17:31:19.299915Z","shell.execute_reply.started":"2022-07-09T17:31:19.241659Z","shell.execute_reply":"2022-07-09T17:31:19.299135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nReads in data from os.\n\"\"\"\ndef set_data_from_path(self, pd=pd, **kwargs):\n  for kwarg in kwargs:\n    assert kwarg.startswith('data_'), \"Passed kwargs must start with 'data_'\"\n    self._attr(**{kwarg:pd.read_csv(kwargs[kwarg])})\n  return self\nAugury.constructor()(set_data_from_path)\naugur.set_data_from_path(data_raw_=TRAIN_PATH)\n\"\"\"\n\"\"\"\n# Output the type of data to get an understanding between continuous and categorical.\naugur.data_raw_.info()","metadata":{"id":"gENcz7_mxGUu","outputId":"48db7a49-cb71-4a29-a57e-10c12c8c8d29","execution":{"iopub.status.busy":"2022-07-09T17:31:19.301035Z","iopub.execute_input":"2022-07-09T17:31:19.301499Z","iopub.status.idle":"2022-07-09T17:31:19.356870Z","shell.execute_reply.started":"2022-07-09T17:31:19.301471Z","shell.execute_reply":"2022-07-09T17:31:19.355746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Explore Data\n\n### Set Label\nReview the label and specify whether the label is `categorical` or `continous`. ","metadata":{"id":"DEWeVmQkj-Ze"}},{"cell_type":"code","source":"LABEL_CONFIG = {\n    \"label\": \"Transported\",       # Column name to use as the label\n    \"label_type\" : \"categorical\"  # [categorical|continuous]\n}\n\"\"\"\n\"\"\"\ndef set_label(self, label=\"\", label_type=\"categorical\"):\n  data_raw, label_types, = self._attr('data_raw_', valid_label_types_=['categorical','continuous'])\n  assert label_type in label_types, label_types\n  self._attr(label_type_=label_type, label_=label, labels_unique_=data_raw[label].unique())\nAugury.constructor()(set_label)\naugur.set_label(**LABEL_CONFIG)\n\"\"\"\n\"\"\"\nprint(\"Label has been set to: \\\"{}\\\", and is a {} target.\".format(augur.label_, augur.label_type_))","metadata":{"id":"WbiNdcgytfyK","outputId":"660f5283-cd21-4154-c45b-362b86968df0","execution":{"iopub.status.busy":"2022-07-09T17:31:19.359571Z","iopub.execute_input":"2022-07-09T17:31:19.360113Z","iopub.status.idle":"2022-07-09T17:31:19.370639Z","shell.execute_reply.started":"2022-07-09T17:31:19.360059Z","shell.execute_reply":"2022-07-09T17:31:19.369771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Categorical Features\nSpecify columns containing categorical data, so that they can be explored for insights.","metadata":{"id":"BhYoWHrynMlZ"}},{"cell_type":"code","source":"# Looking for Dtypes which are not 'float64|int64'\naugur.data_raw_.infer_objects().info()","metadata":{"id":"IAnqYPDP8rxi","outputId":"efa5c0ad-dcbc-43d8-8f29-6de8af5972be","execution":{"iopub.status.busy":"2022-07-09T17:31:19.371802Z","iopub.execute_input":"2022-07-09T17:31:19.372286Z","iopub.status.idle":"2022-07-09T17:31:19.395634Z","shell.execute_reply.started":"2022-07-09T17:31:19.372257Z","shell.execute_reply":"2022-07-09T17:31:19.394526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add headings for continuous features to the following variable. \nCATEGORICAL_FEATURES = [\"PassengerId\", \"HomePlanet\",\"CryoSleep\",\"Cabin\",\"Destination\",\"VIP\", \"Name\"]\n\"\"\"\n\"\"\"\ndef set_data_categorical(self, feature_names):\n  data_raw, label = self._attr('data_raw_', 'label_')\n  self._attr(categorical_features_=feature_names, \n             data_categorical_=data_raw[feature_names + [label]])\nAugury.constructor()(set_data_categorical)\naugur.set_data_categorical(CATEGORICAL_FEATURES)\n\"\"\"\n\"\"\"\n# Summarize the categorical Data by Label to see if any patterns\naugur.data_categorical_.groupby(augur.label_).count()","metadata":{"id":"ftqYrukE8QJT","outputId":"d8a877da-b557-4b5c-cbda-1d83b9b08fc3","execution":{"iopub.status.busy":"2022-07-09T17:31:19.397278Z","iopub.execute_input":"2022-07-09T17:31:19.397960Z","iopub.status.idle":"2022-07-09T17:31:19.425549Z","shell.execute_reply.started":"2022-07-09T17:31:19.397918Z","shell.execute_reply":"2022-07-09T17:31:19.424178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nReport to help identify the categorical columns and see which ones might be\ngood candidates for feature engineering.\n\"\"\"\ndef categorical_unique_report(self):\n  features, data = self._attr('categorical_features_', 'data_categorical_')\n  for feat in features:\n    print('{}: {} unique values'.format(feat, data[feat].nunique()))\nAugury.constructor()(categorical_unique_report)\naugur.categorical_unique_report()\n\"\"\"\n\"\"\"\n# Check for columns with many 'unique' values - could be good candidate for feature engineering.","metadata":{"id":"MuPDk5jPDriI","outputId":"2ff08607-3b7e-408c-fd67-da1fb809ee26","execution":{"iopub.status.busy":"2022-07-09T17:31:19.428467Z","iopub.execute_input":"2022-07-09T17:31:19.428799Z","iopub.status.idle":"2022-07-09T17:31:19.445219Z","shell.execute_reply.started":"2022-07-09T17:31:19.428772Z","shell.execute_reply":"2022-07-09T17:31:19.444306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\nReport to outline the specific values within each categorical columns and their\nfrequency.\n\"\"\"\ndef value_count_report(self):\n  features, data = self._attr('categorical_features_', 'data_categorical_')\n  for feat in features:\n    print(\"\\n======= {} Value Counts =======\".format(feat))\n    print(data[feat].value_counts())\nAugury.constructor()(value_count_report)\naugur.value_count_report()\n\"\"\"\n\"\"\"\n# Get a better sense of the type of data being worked with.","metadata":{"id":"4jhCSPzs-OiM","outputId":"9d5d93fb-c83f-44f6-d544-9bdc38ca40e2","execution":{"iopub.status.busy":"2022-07-09T17:31:19.446505Z","iopub.execute_input":"2022-07-09T17:31:19.446884Z","iopub.status.idle":"2022-07-09T17:31:19.475639Z","shell.execute_reply.started":"2022-07-09T17:31:19.446855Z","shell.execute_reply":"2022-07-09T17:31:19.474538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Handle Nulls and Outliers","metadata":{"id":"H6gXkzDInQ3-"}},{"cell_type":"code","source":"\"\"\"\n  Returns a text-based report of null values, compared against the label.\n  If nothing is returnd than there are no null values.\n\"\"\"\ndef null_report(self, feature_type='continuous'):\n  valid_label_types, label = self._attr('valid_label_types_', 'label_')\n  assert feature_type in valid_label_types, valid_label_types\n  _features, _data = self._attr('%s_features_' % feature_type, \n                                'data_%s_' % feature_type)\n  null_report_ = _data[_features].isnull().sum()\n  print(\"Null Report:\\n===========\")\n  for index, value in null_report_.items():\n    if value == 0:\n      continue\n    print(\"{} is Null (False/True) compared to {} -> Total Nulls:{}\".format(index, label, value))\n    print(\"===========\")\n    print(_data[label].groupby([_data[label],_data[index].isnull()]).count())\n    print(\"===========\")\n  print(\"{} Columns:\\n{}\".format(feature_type.capitalize(),_features))\nAugury.constructor()(null_report)\naugur.null_report(feature_type='categorical')\n\"\"\"\n\"\"\"\n# View Distribution of Nulls Across Label Types","metadata":{"id":"PSh8ta4xAQZr","outputId":"c30d4ec7-760c-4481-d21e-74ef66f206eb","execution":{"iopub.status.busy":"2022-07-09T17:31:19.477178Z","iopub.execute_input":"2022-07-09T17:31:19.477466Z","iopub.status.idle":"2022-07-09T17:31:19.511065Z","shell.execute_reply.started":"2022-07-09T17:31:19.477440Z","shell.execute_reply":"2022-07-09T17:31:19.509900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Based on report seems to be a pretty even split for null rows\n# I'm going to add some 3rd categorical value.\n# Will not include Name\nCATEGORICALS_HANDLE_NULLS = [{\n    \"columns\": [\"HomePlanet\", \"CryoSleep\", \"Destination\", \"VIP\"],\n    \"value\" : \"SomeThirdValue\"\n},{\n    \"columns\": ['Cabin'],\n    \"value\" : \"A/0/A\"\n}]\n\"\"\"\n\"\"\"\ndef best_ref_data(self, *, feature, data_raw, data_clean):\n  \"\"\"\n  Sets a \"cached\" type data structure on the object, so that\n  nulls can be reapplied without causing errors.\n  \"\"\"\n  _suffix = [\"_tr\",\"_clean\"]\n  for _s in _suffix:\n    _column = \"%s%s\" % (feature, _s)\n    if _column in data_clean.columns:\n      return data_clean[_column]\n  return data_raw[feature]\ndef process_nulls(self, data=('data_clean_', 'data_raw_')):\n  null_pipeline, data_raw, label = self._attr('null_pipeline_', \n                                              'data_raw_', 'label_')\n  self._attr(overwrite=False, data_clean_=pd.DataFrame(data_raw[label]), \n               data_test_clean_=pd.DataFrame())\n  _target, _ref = self._attr(*data)\n  for rule in null_pipeline:\n    for column in rule['columns']:\n      column_ = \"%s_clean\" % column\n      if rule['op'] == 'drop':\n        # dropping from _clean will require adding column to pd then dropping\n        _target[column_] = _ref[column]\n        _target.dropna(subset=[column_], inplace=True)\n        # drop inplace for data_raw should update views categorical and continous\n        _ref.dropna(subset=[column], inplace=True)\n      else:\n        _target[column_] = self.best_ref_data(\n            feature=column, data_raw=_ref, data_clean=_target).fillna(rule['value'])\n  return self\ndef add_null_rule(self, *rules):\n  null_pipeline, = self._attr(overwrite=False, null_pipeline_=[])\n  for rule in rules:\n    columns, op, value = self._hash(rule, 'columns', 'op', 'value')\n    data_raw, = self._attr('data_raw_')\n    assert (set(columns) and set(columns).issubset(data_raw.columns)), columns\n    rule['op'] = 'fill' if op is None else op\n    assert rule['op'] in ['fill', 'drop']\n    if rule['op'] == 'fill':\n      assert not value is None, \"Missing value\"\n    null_pipeline.append(rule)\ndef add_null_rules(self, rules, remove_existing=True):\n  if remove_existing and hasattr(self, 'null_pipeline_'):\n    delattr(self, \"null_pipeline_\")\n  self.add_null_rule(*rules)\n  return self\nAugury.constructor()(process_nulls, add_null_rule, add_null_rules, best_ref_data)\naugur.add_null_rules(CATEGORICALS_HANDLE_NULLS).process_nulls()\n\"\"\"\n\"\"\"\n# Review Cleaned Data\naugur.data_clean_","metadata":{"id":"WpShzw_CDPWZ","outputId":"cb070164-8ced-40e6-a279-7d0d8970801c","execution":{"iopub.status.busy":"2022-07-09T17:31:19.515708Z","iopub.execute_input":"2022-07-09T17:31:19.516011Z","iopub.status.idle":"2022-07-09T17:31:19.553569Z","shell.execute_reply.started":"2022-07-09T17:31:19.515985Z","shell.execute_reply":"2022-07-09T17:31:19.552478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Categorical Feature Engineering\nBased on Data Exploration will do the following:\n1. Convert `PassengerId` a group column and id column\n2. Convert `Cabin` to a Deck, Side and number column","metadata":{"id":"zdN6gOysnbk5"}},{"cell_type":"code","source":"FEATURE_ENGINEERING = [{\n  \"column\": \"Cabin\",\n  \"apply\" : lambda x: 'NoCabin' if x == 'NoCabin' else x.split('/')[0],\n  \"target_column\" : \"CabinDeck\"\n},{\n  \"column\": \"Cabin\",\n  \"apply\" : lambda x: 'NoCabin' if x == 'NoCabin' else x.split('/')[2],\n  \"target_column\" : \"CabinSide\"\n},{\n  \"column\": \"Cabin\",\n  \"apply\" : lambda x: -1 if x == 'NoCabin' else int(x.split('/')[1]),\n  \"change_to\": \"continuous\"\n},{\n  \"column\": \"PassengerId\",\n  \"apply\" : lambda x: int(x.split('_')[0]),\n  \"change_to\": \"continuous\",\n  \"target_column\" : \"PassengerGroupId\"\n},{\n  \"column\": \"PassengerId\",\n  \"apply\" : lambda x: int(x.split('_')[1]),\n  \"change_to\": \"continuous\",\n}]\n\"\"\"\nFEATURE_ENGINEERING\n-------\n  List of columnes and callable functions to perform for each row\n\n  Parameters\n  -------\n    column    : str\n      Column to apply the transformation to.\n    apply     : callable(x)\n      Function to appply at each row\n    change_to : (continuous|categorical)\n      Switches the column to the target type specified.\n    target_columns : str\n      Will create a new feature with this name. \n\"\"\"\ndef __data_cache_update_check(self, *, column, data):\n  data_clean, data_raw, data_cache = self._attr(*data, '%scache_' % data[0])\n  _column = \"%s_clean\" % column\n  if _column in data_clean and not _column in data_cache.columns:\n    data_cache[_column] = data_clean[_column]\n    return\n  if not _column in data_cache.columns:\n    data_cache[_column] = data_raw[column]\ndef __swap_features(self,*, change_to, column, target_column):\n  if change_to is None: \n    return\n  valid_changes, = self._attr('valid_label_types_')\n  assert change_to in valid_changes\n  _valid_changes = valid_changes.copy()\n  _to = _valid_changes.pop(_valid_changes.index(change_to))\n  _from = _valid_changes.pop()\n  _feature = column if target_column is None else target_column\n  _to_list, = self._attr(overwrite=False, **{\"%s_features\" % _to : []})\n  try:\n    _from_list = getattr(self, \"%s_features\" % _from)\n    _i = _from_list.index(_feature)\n    _feature = _from_list.pop(_i)\n  except:\n    print(\"Feature may have already transitioned\")\n  _to_list.append(_feature)\n  self._attr(**{\"%s_features\" % _to : list(set(_to_list))})\ndef process_features(self, data=('data_clean_', 'data_raw_')):\n  feature_engineering_pipeline, data_cache, = self._attr('feature_engineering_pipeline_',\n      overwrite=False, **{\"%scache_\" % data[0] : pd.DataFrame()})\n  _target, _ref = self._attr(*data)\n  for rule in feature_engineering_pipeline:\n    column, change_to, apply, target_column, columns = \\\n    self._hash(rule, 'column', 'change_to', 'apply', 'target_column', 'columns')\n    self.__data_cache_update_check(column=column, data=data)\n    self.__swap_features(change_to=change_to, column=column, target_column=target_column)\n    column_ = \"%s_clean\" % (rule['column'] if target_column is None else target_column)\n    _target[column_] = self.best_ref_data(\n        feature=rule['column'], data_raw=_ref, data_clean=data_cache).apply(apply)\n    if target_column is None:\n      _ref[column] = _target[column_]\n    else:\n      _ref[target_column] = _target[column_]\n  return self\ndef __add_feature_engineering_rule(self, *rules):\n  data_raw, feature_engineering_pipeline = self._attr('data_raw_', overwrite=False, \n                         feature_engineering_pipeline_=[])\n  for rule in rules:\n    column, apply, change_to = self._hash(rule, 'column', 'apply', 'change_to')\n    assert (set([column]) and set([column]).issubset(data_raw.columns)), \\\n      \"{} doesn't exists in the original dataset, unable to apply rule.\".format(column)\n    assert callable(apply), \"'apply' must be a callabeble function.\"\n    feature_engineering_pipeline.append(rule)\ndef add_feature_engineering_rules(self, rules, remove_existing=True):\n  if remove_existing and hasattr(self, 'feature_engineering_pipeline_'):\n    delattr(self, \"feature_engineering_pipeline_\")\n  self.__add_feature_engineering_rule(*rules)\n  return self\nAugury.constructor()(add_feature_engineering_rules, __add_feature_engineering_rule,\n                     process_features, __data_cache_update_check, __swap_features)\naugur.add_feature_engineering_rules(FEATURE_ENGINEERING).process_features()\n\"\"\"\n\"\"\"\n# Review recently updated Data\naugur.data_clean_","metadata":{"id":"sCj29jhjUf6Q","outputId":"bf1100f5-f3de-4390-e2ee-2d4f4625465a","execution":{"iopub.status.busy":"2022-07-09T17:31:19.554962Z","iopub.execute_input":"2022-07-09T17:31:19.555778Z","iopub.status.idle":"2022-07-09T17:31:19.627184Z","shell.execute_reply.started":"2022-07-09T17:31:19.555742Z","shell.execute_reply":"2022-07-09T17:31:19.626164Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Continous Data\nSpecify columns containing continous data","metadata":{"id":"5B__ssc-mzLX"}},{"cell_type":"code","source":"# Statistics on the Continous Features\naugur.data_raw_.describe()","metadata":{"id":"mxNS4kgd3OBu","outputId":"87091d91-02b6-42c1-b699-1c135245624a","execution":{"iopub.status.busy":"2022-07-09T17:31:19.628558Z","iopub.execute_input":"2022-07-09T17:31:19.628879Z","iopub.status.idle":"2022-07-09T17:31:19.671118Z","shell.execute_reply.started":"2022-07-09T17:31:19.628851Z","shell.execute_reply":"2022-07-09T17:31:19.669958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add headings for continuous features to the following variable. \nCONTINUOUS_FEATURES = ['PassengerId', 'Cabin', \"Age\", \"RoomService\", \"FoodCourt\",\n                       \"ShoppingMall\", \"Spa\", \"VRDeck\", 'PassengerGroupId']\n\"\"\"\n\"\"\"\ndef set_data_continuous(self, feature_names):\n  data_raw, label= self._attr('data_raw_','label_')\n  self._attr(continuous_features_=feature_names, \n             data_continuous_=data_raw[feature_names + [label]])\nAugury.constructor()(set_data_continuous)\naugur.set_data_continuous(CONTINUOUS_FEATURES)\n\"\"\"\n\"\"\"\n# View Continuous data broken down by Label\naugur.data_continuous_.groupby(augur.label_).mean()","metadata":{"id":"4dB4yxlQ39vf","outputId":"9ccc4747-cd5d-4625-f5e5-c66e3986d08e","execution":{"iopub.status.busy":"2022-07-09T17:31:19.672576Z","iopub.execute_input":"2022-07-09T17:31:19.672960Z","iopub.status.idle":"2022-07-09T17:31:19.697580Z","shell.execute_reply.started":"2022-07-09T17:31:19.672932Z","shell.execute_reply":"2022-07-09T17:31:19.696516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Handle Nulls","metadata":{"id":"iaDHS55lm_aS"}},{"cell_type":"code","source":"\"\"\"\n\"\"\"\n# View Null Continuous Variables by Label\naugur.null_report()","metadata":{"id":"6w58cLyEsHkx","outputId":"190b8e10-7b28-4cd3-8cc6-23dfb7adcfa9","execution":{"iopub.status.busy":"2022-07-09T17:31:19.698751Z","iopub.execute_input":"2022-07-09T17:31:19.699863Z","iopub.status.idle":"2022-07-09T17:31:19.721114Z","shell.execute_reply.started":"2022-07-09T17:31:19.699823Z","shell.execute_reply":"2022-07-09T17:31:19.719965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The following is based on understanding the data from the data set \n# - you may need to tweak for your specific use-case.\n# Based on the above null report there seems to be a pretty even breakdown across\n# all continous variables of null and the target variable, as such:\n# For age we'll use the mean\n# For all others we'll fill with 0s assuming they didn't purchase anything. \nCONTINUOUS_FEATURES_FILL_NULLS = [{\n    \"columns\": [\"Age\"],\n    \"value\" : augur.data_raw_['Age'].mean()\n},{\n    \"columns\": [\"RoomService\", \"FoodCourt\", \"ShoppingMall\", \"Spa\", \"VRDeck\"],\n    \"value\" : 0\n}]\n\"\"\"\nCONTINUOUS_FEATURES_FILL_NULLS\n-------\nArray of Dictionary containing rules to be applied to the dataset.\n\n  Parameters\n  ----------\n  columns : list(str)\n      Column names\n  value   : any\n      Value to replace null with\n  op      : fill | drop\n      fill - will fill any missing nulls\n      drop - will drop any rows with missing null\n\"\"\"\naugur.add_null_rules(\n    CONTINUOUS_FEATURES_FILL_NULLS, remove_existing=False).process_nulls()\n\"\"\"\n\"\"\"\n# View updated Data\naugur.data_clean_","metadata":{"id":"F7KWWtTGseBD","outputId":"eac40521-5423-42a5-8820-3ca085e65a34","execution":{"iopub.status.busy":"2022-07-09T17:31:19.722733Z","iopub.execute_input":"2022-07-09T17:31:19.723167Z","iopub.status.idle":"2022-07-09T17:31:19.762814Z","shell.execute_reply.started":"2022-07-09T17:31:19.723127Z","shell.execute_reply":"2022-07-09T17:31:19.761500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Cap and Floor Outliers","metadata":{"id":"KPDRT-49nDgR"}},{"cell_type":"code","source":"import numpy as np\n\"\"\"\nContinous Outlier Report\nDisplays each continous features and the number of outliers depending on:\n  - 95th percentile\n  - more than 3 standard deviations away from the mean\n  - 99th percentile\n\"\"\"\ndef continuous_outlier_report(self, np=np):\n  continuous_features, data_clean= self._attr('continuous_features_', 'data_clean_')\n  for i, feature in enumerate(continuous_features):\n    column_ = \"%s_clean\" % feature\n    data = data_clean[column_]\n    mean = np.mean(data)\n    std = np.std(data)\n    setattr(self, \"%s_stats_\" % column_,{\n        \"95p\" : data.quantile(.95),\n        \"3sd\" : mean + 3*(std),\n        \"99p\" : data.quantile(.99)\n    })\n    print('\\nOutlier caps for {}:'.format(feature))\n    print('  --95p: {:,.1f} / {} values exceed that'.format(\n      data.quantile(.95),\n      data[data > data.quantile(.95)].shape[0]\n    ))\n    print('  --3sd: {:,.1f} / {} values exceed that'.format(\n      mean + 3*(std), \n      data[data > (mean + 3*(std))].shape[0]\n    ))\n    print('  --99p: {:,.1f} / {} values exceed that'.format(\n      data.quantile(.99),\n      data[data > data.quantile(.99)].shape[0]\n    ))\nAugury.constructor()(continuous_outlier_report)\naugur.continuous_outlier_report()\n\"\"\"\n\"\"\"\n# View Breakdown of Outliers","metadata":{"id":"NHn9judr105D","outputId":"a3436331-11e3-478b-b63f-10c3447b5444","execution":{"iopub.status.busy":"2022-07-09T17:31:19.764198Z","iopub.execute_input":"2022-07-09T17:31:19.765139Z","iopub.status.idle":"2022-07-09T17:31:19.845759Z","shell.execute_reply.started":"2022-07-09T17:31:19.765101Z","shell.execute_reply":"2022-07-09T17:31:19.844532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Based on Above set some rules for clipping the upper limit of the data:\nCONTINUOUS_FEATURES_CLIP = [{\n    \"columns\": [\"Age\", \"PassengerId\"],\n    \"upper\" : \"99p\"\n},{\n    \"columns\": [\"PassengerGroupId\", \"Cabin\", \"RoomService\", \"FoodCourt\", \n                \"ShoppingMall\", \"Spa\", \"VRDeck\"],\n    \"upper\" : \"3sd\"\n}]\n\"\"\"\nCONTINUOUS_FEATURES_CLIP\n-------\n  Sets an upper limit per each column.\n\n  Parameters\n  -------\n    columns : list(str)\n      Columns to be cliped\n    upper   : (95p|99p|3sd)\n      Method to use for clipping\n\"\"\"\ndef clip_continuous(self, data=('data_clean_', 'data_raw_')):\n  clip_pipeline, = self._attr('clip_pipeline_')\n  _target, _ref = self._attr(*data)\n  for rule in clip_pipeline:\n    for column in rule['columns']:\n      _column = \"%s_clean\" % column\n      _upper = getattr(self,\"%s_stats_\" % _column)[rule['upper']]\n      _target[_column] = self.best_ref_data(feature=column, data_raw=_ref, \n                                            data_clean=_target).clip(upper=_upper)\n      print(\"Clipping {} to new max of: {}\".format(_column,_upper))\n  print(\"Updated attr `data_clean`\")\n  return self\ndef __add_clip_rule(self, *rules):\n  clip_pipeline, = self._attr(overwrite=False, clip_pipeline_=[])\n  for rule in rules:\n    columns, upper = self._hash(rule, 'columns', 'upper')\n    data_raw, =self._attr('data_raw_')\n    if set(columns) and not set(columns).issubset(data_raw.columns):\n      return print(\"{} doesn't exists in the original dataset, unable to apply rule.\".format(columns))\n    valid_upper_args = ['3sd', '95p', '99p']\n    assert upper in valid_upper_args, \"'upper' must be one of:%s\" % valid_upper_args\n    clip_pipeline.append(rule)\ndef add_clip_rules(self, rules, remove_existing=False):\n  if remove_existing and hasattr(self, 'clip_pipeline_'):\n    delattr(self, \"clip_pipeline_\")\n  self.__add_clip_rule(*rules)\n  return self\nAugury.constructor()(clip_continuous, __add_clip_rule, add_clip_rules)\naugur.add_clip_rules(CONTINUOUS_FEATURES_CLIP, remove_existing=True).clip_continuous()\n\"\"\"\n\"\"\"\naugur.data_clean_.describe()","metadata":{"id":"ZoPQ-bjp4XG7","outputId":"0197db17-7fba-4d81-c6a2-d76e332d6c9d","execution":{"iopub.status.busy":"2022-07-09T17:31:19.847504Z","iopub.execute_input":"2022-07-09T17:31:19.848269Z","iopub.status.idle":"2022-07-09T17:31:19.905869Z","shell.execute_reply.started":"2022-07-09T17:31:19.848226Z","shell.execute_reply":"2022-07-09T17:31:19.905145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Continuous Feature Engineering","metadata":{"id":"tpTuE_6aMwZ5"}},{"cell_type":"code","source":"CONTINUOUS_FEATURE_ENGINEERING = [{\n    \"columns\": [\"RoomService\", \"FoodCourt\",\"ShoppingMall\", \"Spa\", \"VRDeck\"],\n    \"apply\" : lambda x: x.sum(),\n    \"target_column\" : \"TotalBill\"\n},{\n    \"columns\": [\"RoomService\", \"FoodCourt\",\"ShoppingMall\", \"Spa\", \"VRDeck\"],\n    \"apply\" : lambda x: x.count(),\n    \"condition\" : lambda columns: columns > 1,\n    \"target_column\" : \"TotalServices\"\n}]\n\"\"\"\nCONTINUOUS_FEATURE_ENGINEERING\n-------\n  List of columnes and callable functions to perform for each row.\n\n  Parameters\n  -------\n    columns    : list(str)\n      Column to apply the transformation to.\n    apply     : callable(x)\n      Function to appply at each row\n    target_column : str\n      New Column to be added to the dataset.\n    condition : callable(x)\n      Optional, function to be applied to all columns prior to calling 'apply'\n\"\"\"\ndef __add_to_feature_list(self,is_a='continuous',*, feature):\n  valid_changes, = self._attr('valid_label_types_')\n  assert is_a in valid_changes\n  _valid_changes = valid_changes.copy()\n  _to = _valid_changes.pop(_valid_changes.index(is_a))\n  _to_list, = self._attr(\"%s_features_\" % _to)\n  _to_list.append(feature)\n  self._attr(**{\"%s_features_\" % _to : list(set(_to_list))})\ndef __call_apply(sef, *, data, columns, condition, apply):\n  _condition = columns if condition is None else condition(data[columns]) \n  _data = data[columns]\n  return _data[_condition].apply(apply, axis=1)\ndef process_continuous_features(self, data=('data_clean_', 'data_raw_')):\n  continuous_feature_engineering_pipeline, = \\\n    self._attr('continuous_feature_engineering_pipeline_')\n  _target, _ref = self._attr(*data)\n  for rule in continuous_feature_engineering_pipeline:\n    columns, apply, target_column, condition = \\\n    self._hash(rule, 'columns', 'apply', 'target_column', 'condition')\n    columns_ = [\"%s_clean\" % c for c in columns] \n    # Adding to both datasets\n    _ref[target_column] = self.__call_apply(data=_ref, columns=columns, \n                                            condition=condition, apply=apply)\n    _target[\"%s_clean\" % target_column] = self.__call_apply(data=_target, columns=columns_, \n                                               condition=condition, apply=apply)\n    self.__add_to_feature_list(feature=target_column)\n  return self\ndef __add_continuous_feature_engineering_rule(self, *rules):\n  continuous_feature_engineering_pipeline, = self._attr(overwrite=False, \n                         continuous_feature_engineering_pipeline_=[])\n  for rule in rules:\n    columns, apply, target_column, condition = \\\n    self._hash(rule, 'columns', 'apply', 'target_column', 'condition')\n    assert callable(apply), \"'apply' must be a callable function.\"\n    rule['condition'] = None if not callable(condition) else condition\n    continuous_feature_engineering_pipeline.append(rule)\ndef add_continuous_feature_engineering_rules(self, rules, remove_existing=True):\n  if remove_existing and hasattr(self, 'continuous_feature_engineering_pipeline_'):\n    delattr(self, \"continuous_feature_engineering_pipeline_\")\n  self.__add_continuous_feature_engineering_rule(*rules)\n  return self\nAugury.constructor()(add_continuous_feature_engineering_rules, \n                     process_continuous_features,\n                     __add_continuous_feature_engineering_rule, __call_apply, \n                     __add_to_feature_list)\naugur.add_continuous_feature_engineering_rules(\n    CONTINUOUS_FEATURE_ENGINEERING).process_continuous_features()\n\"\"\"\n\"\"\"\naugur.data_clean_","metadata":{"id":"UAVPkApN3QI7","outputId":"7a070c27-977c-4c4d-b156-5e4bfa49ae89","execution":{"iopub.status.busy":"2022-07-09T17:31:19.906993Z","iopub.execute_input":"2022-07-09T17:31:19.907464Z","iopub.status.idle":"2022-07-09T17:31:21.175586Z","shell.execute_reply.started":"2022-07-09T17:31:19.907436Z","shell.execute_reply":"2022-07-09T17:31:21.174535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Transform Skewed Features\nThis is probably the most manually invovled section, since each column needs to be reviewed for shape and transformation.\n\nAs a first step will plot each feature in relation to the target variable to get an idea of shapes each feature represents.","metadata":{"id":"6OXOvlaAnFn0"}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nfrom datetime import datetime as dt\n%matplotlib inline\n\"\"\"\nWill generate a report with 3 columns, and as many rows as there are features\nto fill those columns representing the data in comparison to the label data.\n\nThe goal of this function is to help visualize the data. \n\"\"\"\ndef continuous_shape_report(self, plt=plt, dt=dt, sns=sns, np=np):\n  continuous_features, data_clean, label, labels_unique = \\\n    self._attr('continuous_features_', 'data_clean_', 'label_', 'labels_unique_')\n  num_rows = len(continuous_features) // 3 + (len(continuous_features) % 3 > 0)\n  fig,axes = plt.subplots(num_rows,3, figsize=(15,5 * num_rows))\n  prop_cycle = plt.rcParams['axes.prop_cycle']\n  colors = prop_cycle.by_key()['color']\n  colors_len = len(colors)\n  row_i, col_i = 0, 0\n  start_ = dt.now()\n  for feature in continuous_features:\n    column_ = \"%s_clean\" % feature\n    bars = [data_clean[data_clean[label] == _label][column_] \n                 for _label in labels_unique]\n    xmin = min([_bar.min() for _bar in bars])\n    xmax = max([_bar.max() for _bar in bars])\n    width = (xmax - xmin) / 40\n    ax = axes[row_i,col_i]\n    for _i, _bar in enumerate(bars):\n      sns.histplot(list(_bar), color=colors[_i % len(colors)],bins=np.arange(xmin,xmax,width), ax=ax)\n    ax.legend(labels_unique)\n    ax.title.set_text('Overlaid histogram for {}'.format(feature))\n    col_i = (col_i + 1) % 3\n    row_i = row_i + 1 if col_i % 3 == 0 else row_i\n  # Clean up any empty columns\n  while not col_i % 3 == 0:\n    axes[row_i,col_i].set_axis_off()\n    col_i += 1\n  #plt.tight_layout()\n  plt.show()\n  # Tracking latency since it's a bit of long running function.\n  print(\"Latency (hh:mm:ss.ms): {}\".format(dt.now() - start_))\n  print(\"Features mapped: {}\".format(continuous_features))\nAugury.constructor()(continuous_shape_report)\naugur.continuous_shape_report()\n\"\"\"\n\"\"\"\n# Graphs to review shape of continuous features","metadata":{"id":"laWiCaH6Pc2W","outputId":"8116cdbb-f713-42d6-dd81-a13bd3ee8740","execution":{"iopub.status.busy":"2022-07-09T17:31:21.177963Z","iopub.execute_input":"2022-07-09T17:31:21.178456Z","iopub.status.idle":"2022-07-09T17:31:25.337786Z","shell.execute_reply.started":"2022-07-09T17:31:21.178411Z","shell.execute_reply":"2022-07-09T17:31:25.337001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Based on the above plots - we may need to transform some of them so that they appear more Gaussian in Distribution.","metadata":{"id":"1GVfGav99QuL"}},{"cell_type":"code","source":"from scipy.stats import boxcox\n### BoxCox Plots\n# For the specific examples from `titanic-spaceship` we can see that Age is normally distributed\n# but we'll include it to confirm the shape remains consistent - however the other features which\n# relate to purchases seem to have more of a poisson type curve, where there\n# are a lot of individuals who did not make any purchases, as such it might be\n# beneficial to transform these to more normal looking shapes.\nCONTINUOUS_FEATURES_BOXCOXPLOTS = {\n    \"columns\": [\"RoomService\", \"FoodCourt\", \"ShoppingMall\", \"Spa\", \"VRDeck\", \"Age\",\n                \"Cabin\", \"PassengerGroupId\", \"PassengerId\", \"TotalBill\", \"TotalServices\"],\n}\n\"\"\"\nPlot Box Cox transformations, adds \"_tr\" columns to the data_clean attribute, \nand records the lambda used in the \"_stats\" attr.\n\"\"\"\ndef __call_boxcox(self, lmbda=None, boxcox=boxcox, *, data, column):\n  c = abs(data[column].min()) + 1\n  _data = data[column] + c\n  return boxcox(_data) if lmbda is None else boxcox(_data, lmbda)  \ndef __box_cox_plots(self, *features, plt=plt, dt=dt, sns=sns):\n  data_clean, = self._attr('data_clean_')\n  _rows = len(features) // 2 + (len(features) % 2 > 0)\n  fig,axes = plt.subplots(_rows, 4, figsize=(5 * _rows, 5 * _rows))\n  row_i, col_i = 0, 0\n  start_ = dt.now()\n  transformations = []\n  for feature in features:\n    _column = \"%s_clean\" % feature\n    column_tr = \"%s_tr\" % _column\n    data_clean[column_tr], fitted_lambda = self.__call_boxcox(data=data_clean, \n                                                              column=_column)\n    # Save Lambda\n    transformations.append((feature, fitted_lambda))\n    ax = axes[row_i, col_i]\n    ax.title.set_text(\"{}: Fitted with Lambda:{}\".format(feature, round(fitted_lambda,2)))\n    axes[row_i, col_i + 1].title.set_text(\n        \"{}: Original \".format(feature))\n    sns.kdeplot(data_clean[_column], fill=True, linewidth=2, ax = axes[row_i, col_i + 1])\n    sns.kdeplot(data_clean[column_tr], fill=True, linewidth=2, ax = ax)\n    col_i = (col_i + 2) % 4\n    row_i = row_i + 1 if col_i % 4 == 0 else row_i\n  self._attr(boxcox_pipeline_=transformations)\n  while(col_i % 4 > 0):\n    axes[row_i,col_i].set_axis_off()\n    col_i +=1\n  plt.tight_layout()\n  plt.show()\n  print(\"Latency (hh:mm:ss.ms):{}\".format(dt.now() - start_))\n\ndef box_cox_plots(self, columns=[]):\n  data_raw, = self._attr('data_raw_')\n  assert set(columns) and set(columns).issubset(data_raw.columns), \\\n          \"Columns: {}, do not exists in: {}\".format(columns, data_raw.columns)\n  self.__box_cox_plots(*columns)\nAugury.constructor()(box_cox_plots, __box_cox_plots, __call_boxcox)\naugur.box_cox_plots(**CONTINUOUS_FEATURES_BOXCOXPLOTS)\n\"\"\"\n\"\"\"\n# Review plots of BoxCox Transformations. ","metadata":{"id":"9AjnOXB45Shd","outputId":"f4e53392-342a-42e5-d423-29bece213b5f","execution":{"iopub.status.busy":"2022-07-09T17:31:25.339329Z","iopub.execute_input":"2022-07-09T17:31:25.339646Z","iopub.status.idle":"2022-07-09T17:31:30.012022Z","shell.execute_reply.started":"2022-07-09T17:31:25.339617Z","shell.execute_reply":"2022-07-09T17:31:30.010756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Transform to Numerical Indicators","metadata":{"id":"vurnyaL7nXP0"}},{"cell_type":"code","source":"from sklearn.preprocessing import OrdinalEncoder\nfrom pandas.api.types import is_numeric_dtype\n\n\"\"\"\nChecks data for non-numeric values and converts to integers.\n\"\"\"\ndef __transform(self,*, _target, column, stats_attr):\n  stats_,  = self._attr(overwrite=False, **{stats_attr: {}})\n  try:\n    _target[\"%s_tr\" % column] = stats_['enc'].transform(\n        _target[column].values.reshape(-1,1))\n  except:\n    # Typically thrown due to bool mixed with str... so casting\n    _target[\"%s_tr\" % column] = stats_['enc'].transform(\n        _target[column].astype(str).values.reshape(-1,1))\ndef __encode(self, *, enc, _target, column, stats_attr):\n  stats_,  = self._attr(overwrite=False, **{stats_attr: {}})\n  # Reshaping so that we can encode one column at a time\n  try:\n    stats_['enc'] = enc.fit(_target[[column]].values.reshape(-1,1))\n  except:\n    # Most likely getting an error for booleans where added a 3rd cat.\n    # for handling NaNs\n    stats_['enc'] = enc.fit(_target[[column]].astype(str).values.reshape(-1,1))\n  self._attr(**{stats_attr:stats_})\ndef numeric_transformation(self, data='data_clean_', is_numeric_dtype=is_numeric_dtype,\n                           OrdinalEncoder=OrdinalEncoder, transform=True, encode=True):\n  _target, = self._attr(data)\n  for column in _target.columns:\n    if is_numeric_dtype(_target[column]):\n      continue\n    if column.endswith(\"_tr\"):\n      continue\n    enc = OrdinalEncoder()\n    # Save Encoder (needed for test)\n    stats_attr = \"%s_stats_\" % column\n    if encode == True:\n      self.__encode(enc=enc, _target=_target, column=column, stats_attr=stats_attr)\n    if transform == True:\n      self.__transform(_target=_target, column=column, stats_attr=stats_attr)\n  return self\nAugury.constructor()(numeric_transformation, __transform, __encode)\naugur.numeric_transformation()\n\"\"\"\n\"\"\"\naugur.data_clean_","metadata":{"id":"_AykR5WsiCNm","outputId":"68c75cc5-1e4d-489a-9b0b-bcba4cb0f030","execution":{"iopub.status.busy":"2022-07-09T17:31:30.013811Z","iopub.execute_input":"2022-07-09T17:31:30.014246Z","iopub.status.idle":"2022-07-09T17:31:30.171225Z","shell.execute_reply.started":"2022-07-09T17:31:30.014210Z","shell.execute_reply":"2022-07-09T17:31:30.170507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Feature Scaling & Optimal Count","metadata":{"id":"En6N3fS5nfLw"}},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler, StandardScaler\n\"\"\"\nDisplays Graphs for Comparing MinMax vs Standard.\n\"\"\"\ndef set_data_final(self, test=False):\n  data_clean, label, data_test_clean = self._attr('data_clean_', 'label_', 'data_test_clean_')\n  data_final, data_test_final, = self._attr(\n                  overwrite=False, data_final_=pd.DataFrame(data_clean[label]), \n                  data_test_final_=pd.DataFrame())  \n  _target, _ref = (data_final, data_clean) if test == False else \\\n    (data_test_final, data_test_clean)\n  columns_priority = [_c for _c in data_clean.columns if _c.endswith(\"_tr\")]\n  columns_processed = {}\n  for column_ in columns_priority:\n    column = column_.split(\"_\")[0]\n    _target[column] = _ref[column_]\n    columns_processed[column] = 1\n  for column_ in _ref.columns:\n    column = column_.split(\"_\")[0]\n    if column in columns_processed:\n      continue\n    _target[column] = _ref[column_]\n  return self\ndef scaler_report(self, dt=dt, plt=plt,MinMaxScaler=MinMaxScaler, \n                  StandardScaler=StandardScaler, pd=pd):\n  data_final, label = self._attr('data_final_','label_')\n  fig,axes = plt.subplots(2, 2, figsize=(20, 15))\n  row_i, col_i = 0, 0\n  start_ = dt.now()\n  _features = data_final.drop([label], axis=1)\n  # MinMax\n  minmax_scaler, standard_scaler = self._attr(minmax_scaler_=MinMaxScaler().fit(_features), \n                                              standard_scaler_=StandardScaler().fit(_features))\n  minmax_data, standard_data = \\\n  self._attr(data_minmax_scaled_=pd.DataFrame(minmax_scaler.transform(_features), \n                                   index=_features.index, columns=_features.columns),\n             data_standard_scaled_=pd.DataFrame(standard_scaler.transform(_features),\n                                   index=_features.index, columns=_features.columns))\n  axes[row_i, col_i].title.set_text(\"MinMax Scaler Results\")\n  for _data in [minmax_data, standard_data]:\n    _data.plot(kind='kde', ax=axes[row_i,col_i])\n    _data.plot(kind='hist', bins=20, ax=axes[row_i,col_i + 1])\n    row_i += 1\n  axes[row_i - 1, col_i].title.set_text(\"Standard Scaler Results\")\n  plt.show()\n  print(\"Latency (hh:mm:ss.ms):{}\".format(dt.now() - start_))\nAugury.constructor()(scaler_report, set_data_final)\naugur.set_data_final().scaler_report()\n\"\"\"\n\"\"\"\n# View Results Scaled to decide which scaler to use.","metadata":{"id":"9Q8m2yE91ENA","outputId":"61fa16c1-e0b7-4429-ffb5-b73d34ef54de","execution":{"iopub.status.busy":"2022-07-09T17:31:30.172458Z","iopub.execute_input":"2022-07-09T17:31:30.172982Z","iopub.status.idle":"2022-07-09T17:31:36.632540Z","shell.execute_reply.started":"2022-07-09T17:31:30.172953Z","shell.execute_reply":"2022-07-09T17:31:36.631447Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Based on the above graphs, and knowing that the data generally isn't normally\n# distributed - will set the scaler to be MinMax\nSCALAR_CONFIG = {\n    'scaler' : 'minmax' # minmax | standard\n}\n\"\"\"\nSets the scaler to be used on test.\n\"\"\"\ndef set_scaler(self, scaler='minmax'):\n  valid_scalers = ['minmax', 'standard']\n  assert scaler in valid_scalers, \"Expecting 'scaler' to be one of %s\" % valid_scalers\n  scaler, data = self._attr(\"%s_scaler_\" % scaler, \"data_%s_scaled_\" % scaler)\n  self._attr(scaler_=scaler, data_scaled_=data)\nAugury.constructor()(set_scaler)\naugur.set_scaler(**SCALAR_CONFIG)\n\"\"\"\n\"\"\"\naugur.data_scaled_","metadata":{"id":"BH4Bj_3PL6s6","outputId":"6705ecc1-9987-4a31-faa8-3726c713b0e9","execution":{"iopub.status.busy":"2022-07-09T17:31:36.633988Z","iopub.execute_input":"2022-07-09T17:31:36.634660Z","iopub.status.idle":"2022-07-09T17:31:36.666778Z","shell.execute_reply.started":"2022-07-09T17:31:36.634618Z","shell.execute_reply":"2022-07-09T17:31:36.665810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.feature_selection import RFECV\nimport joblib\n### Feature Analysis using RFECV\n\n\"\"\"\nReport which helps identify key features.\nCould be used to optimize model training.\n\"\"\"\ndef __set_and_fit_rfecv(self, RFECV=RFECV, *, estimator, X, y):\n  _rfecv = RFECV(estimator=estimator)\n  _rfecv.fit(X, y)\n  return _rfecv\ndef __set_compressed_features(self, *, rfecv, data_scaled):\n  features_compressed = []\n  print(\"Features passed to data_compressed_:\")\n  for _i, _f in enumerate(rfecv.ranking_):\n    if _f == 1:\n      feature = rfecv.feature_names_in_[_i]\n      features_compressed.append(feature)\n      print(feature)\n  self._attr(data_compressed_=data_scaled[features_compressed])\ndef feature_analysis(self, max_depth=None, memory=joblib.Memory(AUGURY_CACHE, verbose=0),\n                     RandomForestClassifier=RandomForestClassifier, **kwargs):\n  data_scaled, data_final, label = self._attr('data_scaled_', 'data_final_', 'label_')\n  max_depth = len(data_scaled.columns) // 2 if max_depth is None else max_depth\n  rf = RandomForestClassifier(max_depth=max_depth, **kwargs)\n  start_ = dt.now()\n  cache_rfecv = memory.cache(self.__set_and_fit_rfecv, ignore=['self', 'RFECV'])\n  rfecv = cache_rfecv(RFECV=RFECV, estimator=rf, X=data_scaled, y=data_final[label])\n  time_taken = dt.now() - start_\n  # Plot number of features VS. cross-validation scores\n  plt.figure(figsize=(10,5))\n  plt.xlabel(\"Number of Features Selected\")\n  plt.ylabel(\"Cross validation mean score (nb of correct classifications)\")\n  plt.plot(range(1, len(rfecv.cv_results_['mean_test_score']) + 1), rfecv.cv_results_['mean_test_score'])\n  plt.title(\"Optimal number of features : %d\" % rfecv.n_features_)\n  plt.show()\n  print(\"Latency (hh:mm:ss.ms): {}\".format(time_taken))\n  self.__set_compressed_features(rfecv=rfecv, data_scaled=data_scaled)\nAugury.constructor()(feature_analysis, __set_and_fit_rfecv, __set_compressed_features)\naugur.feature_analysis()\n\"\"\"\n\"\"\"\n# Gut check on fetures to send to model\naugur.data_compressed_","metadata":{"id":"mQnUZPGYOt8N","outputId":"00abab9c-cc90-43a7-fb7c-080904c83a22","execution":{"iopub.status.busy":"2022-07-09T17:31:36.668000Z","iopub.execute_input":"2022-07-09T17:31:36.668344Z","iopub.status.idle":"2022-07-09T17:32:42.845249Z","shell.execute_reply.started":"2022-07-09T17:31:36.668316Z","shell.execute_reply":"2022-07-09T17:32:42.844165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### TODO: Add Correlation Heatmap and process functions for removing functions (using compressed data)","metadata":{"id":"uZg-ifylsJN5","execution":{"iopub.status.busy":"2022-07-09T17:32:42.846796Z","iopub.execute_input":"2022-07-09T17:32:42.847504Z","iopub.status.idle":"2022-07-09T17:32:42.852901Z","shell.execute_reply.started":"2022-07-09T17:32:42.847463Z","shell.execute_reply":"2022-07-09T17:32:42.851930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train / Validation Split","metadata":{"id":"5MutLM_NnlA6"}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\"\"\"\n\"\"\"\ndef confirm_splits(self):\n  # Confirm split\n  for _dataset in [\"y\", \"y_compressed\"]:\n    data_final, label = self._attr('data_final_', 'label_')\n    for _split in [\"%s_train_\" % _dataset, \"%s_val_\" % _dataset]:\n      print(\"%s as %% of rows\" % _split)\n      print(round(len(getattr(self, _split)) / data_final[label].shape[0], 2))\ndef split_hash(self, prefix=\"\",*, X, y):\n  split_tuple = train_test_split(X, y, test_size=0.2)\n  return {\n    'X_%strain_' % prefix  : split_tuple[0], \n    'X_%sval_' % prefix    : split_tuple[1], \n    'y_%strain_' % prefix  : split_tuple[2], \n    'y_%sval_' % prefix    : split_tuple[3]    \n  }\ndef set_train_val_data(self, train_test_split=train_test_split):\n  data_final, data_scaled, data_compressed, label = \\\n    self._attr('data_final_', 'data_scaled_', 'data_compressed_', 'label_')\n  _labels, _features, features_compressed = data_final[label], data_scaled, data_compressed\n  _split_dict = {\n    **self.split_hash(X=_features, y=_labels),\n    **self.split_hash(prefix=\"compressed_\", X=features_compressed, y=_labels),   \n  }\n  self._attr(**_split_dict)\n  return self\nAugury.constructor()(set_train_val_data, confirm_splits, split_hash)\naugur.set_train_val_data().confirm_splits()\n\"\"\"\n\"\"\"\n# Split and then confirm splits","metadata":{"id":"TOeJP-h1YpgJ","outputId":"c3b6f0f4-61ac-469b-8c3d-02e531f0c744","execution":{"iopub.status.busy":"2022-07-09T17:32:42.854340Z","iopub.execute_input":"2022-07-09T17:32:42.854968Z","iopub.status.idle":"2022-07-09T17:32:42.874651Z","shell.execute_reply.started":"2022-07-09T17:32:42.854927Z","shell.execute_reply":"2022-07-09T17:32:42.873770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Principal Component Analysis","metadata":{"id":"2DkVaR8EnieR"}},{"cell_type":"code","source":"from sklearn.decomposition import PCA\n# Perform with data_scaled and continuous features - fed to KMeans\n\"\"\"\n\"\"\"\ndef get_continuous_data(self, X, np=np, dt=dt):\n  continuous_features, = self._attr('continuous_features_')\n  _columns = np.intersect1d(continuous_features, X.columns)\n  return (X[_columns], _columns)\ndef pca_analysis(self, PCA=PCA, data=('X_train_','y_train_','X_val_','y_val_')):\n  X_train, y_train, X_val, y_val = self._attr(*data)\n  X_continuous_train, _ = self.get_continuous_data(X_train)\n  X_continuous_val, _ = self.get_continuous_data(X_val)\n  _explained_variance_ratio, _attempts = 0, 1\n  while(_explained_variance_ratio < 0.80 and _attempts < len(X_train.columns)):\n    _attempts += 1\n    pca=PCA(n_components=(_attempts))\n    pca.fit(X_continuous_train)\n    _explained_variance_ratio = sum(pca.explained_variance_ratio_)\n  start_ = dt.now()\n  if _attempts < 2:\n    # Could not perform PCA\n    X_continuous_train['PCA'] = 1\n    X_continuous_val['PCA'] = 1\n    self._attr(X_pca_train_=X_continuous_train.to_numpy(), y_pca_train_=y_train, \n               X_pca_val_=X_continuous_val.to_numpy(), y_pca_val_=y_val)\n    print(\"Unable to perform PCA\")\n  else:\n    self._attr(X_pca_train_=pca.transform(X_continuous_train), y_pca_train_=y_train)\n    self._attr(X_pca_val_=pca.transform(X_continuous_val), y_pca_val_=y_val)\n    print(\"Latency (hh:mm:ss:ms):{}\".format(dt.now() - start_))\n    print(\"PCA Analysis set {} components. With Total Explained Variance Ratio: {}%\".format(\n        _attempts, round(_explained_variance_ratio * 100,2)))\nAugury.constructor()(pca_analysis, get_continuous_data)\naugur.pca_analysis(data=('X_compressed_train_', 'y_compressed_train_', \n                         'X_compressed_val_', 'y_compressed_val_'))\n\"\"\"\n\"\"\"\n# Rought PCA Analysis (will be used to visualize clusters)","metadata":{"id":"zpAd6eFUWKUI","outputId":"c12b1ce0-0bf6-476f-deb3-595bc71cc05e","execution":{"iopub.status.busy":"2022-07-09T17:32:42.876096Z","iopub.execute_input":"2022-07-09T17:32:42.876438Z","iopub.status.idle":"2022-07-09T17:32:42.907351Z","shell.execute_reply.started":"2022-07-09T17:32:42.876410Z","shell.execute_reply":"2022-07-09T17:32:42.906659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### KMeans Benchmark","metadata":{"id":"c6YoQeS3Y-3q"}},{"cell_type":"code","source":"from sklearn.cluster import KMeans\n# Trying to increase my knowledge of clustering so adding here since it's\n# handled differently than the other models which will be presented.\n\n\"\"\"\nGenerates an Elbow Report for KMeans Analysis\n\"\"\"\ndef kmeans_elbow_report(self, dt=dt, plt=plt):\n  X_pca_train, = self._attr('X_pca_train_')\n  Sum_of_squared_distances = []\n  K = range(2,15)\n  start_ = dt.now()\n  for num_clusters in K :\n    kmeans = KMeans(n_clusters=num_clusters)\n    kmeans.fit(X_pca_train)\n    Sum_of_squared_distances.append(kmeans.inertia_)\n  plt.plot(K,Sum_of_squared_distances,'bx-')\n  plt.xlabel('Values of K') \n  plt.ylabel('Sum of squared error') \n  plt.title('KMeans Clustering for K=2 to K=15')\n  plt.show()\n  print(\"Latency (hh:mm:ss.ms):{}\".format(dt.now() - start_))\nAugury.constructor()(kmeans_elbow_report)\naugur.kmeans_elbow_report()\n\"\"\"\n\"\"\"\n# Elbow report for KMeans Clustering","metadata":{"id":"XoLS58EpNg2K","outputId":"86906396-318f-485d-c8de-007f2b19771f","execution":{"iopub.status.busy":"2022-07-09T17:32:42.911716Z","iopub.execute_input":"2022-07-09T17:32:42.912286Z","iopub.status.idle":"2022-07-09T17:32:58.298542Z","shell.execute_reply.started":"2022-07-09T17:32:42.912253Z","shell.execute_reply":"2022-07-09T17:32:58.297483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\nfrom sklearn.metrics.cluster import contingency_matrix\nKMEANS_ANALYSIS_CONFIG = {\n    'n_clusters' : 5\n}\n\"\"\"\nGenerates a Cluster Report based on PCA-1 and PCA-2\nRelabels according to label based on a contingency matrix, whic is used\nfor traditional accuracy, precision, recal, reporting.\n\"\"\"\ndef print_performance(self, *, model, accuracy, precision, recall, latency):\n  (accuracy, precision, recall) = (round(accuracy * 100, 2), round(precision * 100, 2), round(recall * 100, 2))\n  print(\"{} Performance -- Accuracy: {}% / Precision: {}% / Recall: {}% / Latency (hh:mm:ss.ms): {}\".format(\n      model, accuracy, precision, recall, latency\n  ))\ndef set_performance(self, *, model, accuracy, precision, recall, latency, \n                    output_string=False):\n  models, best_accuracy_, best_precision_, best_recall_, best_latency = self._attr(\n      overwrite=False, models_={}, best_accuracy_=(\"\",0), best_precision_=(\"\",0),\n      best_recall_=(\"\",0), best_latency_=(\"\",0)\n  )\n  _t = { 'accracy' : accuracy, 'precision': precision, 'recall' : recall, \n        'latency': latency } \n  models[model] = {**_t, **models[model]} if model in models else _t\n  self._attr(models_=models)\n  for _top in ['accuracy', \"precision\", 'recall']:\n    _best = \"best_%s_\" % _top\n    if locals()[_top] > locals()[_best][1]:\n      self._attr(**{_best:(model,locals()[_top])})\n  if not best_latency[1] or latency < best_latency[1]:\n    self._attr(best_latency_=(model, latency))\n  if output_string:\n    self.print_performance(model=model, accuracy=accuracy,\n                      precision=precision, recall=recall, latency=latency)\ndef kmeans_plot_axes(self, *, axes, x, y, c, title, plt=plt, np=np):\n  prop_cycle = plt.rcParams['axes.prop_cycle']\n  colors = prop_cycle.by_key()['color']\n  colors_len = len(colors)\n  color_theme = np.array(colors)\n\n  axes.scatter(x=x, y=y, c=color_theme[c % colors_len], s=25)\n  axes.title.set_text(title)\ndef kmeans_results(self, classification_report=classification_report, *, relabel,\n                   time_taken):\n  y_pca_val, = self._attr('y_pca_val_')\n  kmeans_results = classification_report(y_pca_val, relabel, output_dict=True, \n                                         zero_division=0)\n  self.set_performance(model='kmeans', accuracy=kmeans_results['accuracy'],\n                       precision=kmeans_results['weighted avg']['precision'], \n                       recall=kmeans_results['weighted avg']['recall'],\n                       latency=time_taken, output_string=True)\ndef kmeans_analysis(self, n_clusters=2, plt=plt, KMeans=KMeans, dt=dt, np=np, \n                    contingency_matrix=contingency_matrix):\n  X_pca_val, y_pca_val = self._attr('X_pca_val_', 'y_pca_val_')\n  fig,axes = plt.subplots(1,3, figsize=(20,10))\n\n  clustering = KMeans(n_clusters=n_clusters)\n  start_ = dt.now()\n  clustering.fit(X_pca_val)\n  time_taken = dt.now() - start_\n\n  self.kmeans_plot_axes(axes=axes[0], x=X_pca_val[:,0], y=X_pca_val[:,1], \n                    c=y_pca_val, title='Ground Truth Classification')\n  self.kmeans_plot_axes(axes=axes[1], x=X_pca_val[:,0], y=X_pca_val[:,1], \n                    c=clustering.labels_, title='K-Means Classification')\n\n  _n = contingency_matrix(y_pca_val, clustering.labels_)\n  choose_arr = _n.argmax(axis=0)\n  relabel = np.choose(clustering.labels_, choose_arr).astype(np.int64)\n\n  self.kmeans_plot_axes(axes=axes[2], x=X_pca_val[:,0], y=X_pca_val[:,1], \n                    c=relabel, title='K-Means Relabelled Classification')  \n  plt.show()\n  print(\"Latency (hh:mm:ss.ms):{}\".format(time_taken))\n  self.kmeans_results(relabel=relabel, time_taken=time_taken)\nAugury.constructor()(kmeans_analysis,set_performance,kmeans_results, \n                     kmeans_plot_axes, print_performance)\naugur.kmeans_analysis(**KMEANS_ANALYSIS_CONFIG)\n\"\"\"\n\"\"\"\n# Display cluster labels vs. Truth Labels","metadata":{"id":"4Yv42JPQQaGo","outputId":"68c0adfa-b9de-489d-cad1-5a10ca2c436f","execution":{"iopub.status.busy":"2022-07-09T17:32:58.300100Z","iopub.execute_input":"2022-07-09T17:32:58.300412Z","iopub.status.idle":"2022-07-09T17:32:58.859612Z","shell.execute_reply.started":"2022-07-09T17:32:58.300387Z","shell.execute_reply":"2022-07-09T17:32:58.858549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Selection\n|                          | Label Type | Training Speed | Performance Speed | Simplicity | Performance | Performance With Limited Data |\n|:------------------------:|:----------:|:--------------:|:-----------------:|:----------:|:-----------:|:-----------------------------:|\n|    logistic-regression   |    class   |      high      |        high       |     med    |     low     |              high             |\n| support-vector-machines  |    class   |       low      |        med        |     low    |     med     |              high             |\n|   multilayer-perceptron  |    both    |       low      |        med        |     low    |     high    |              low              |\n|       random-forest      |    both    |       med      |        med        |     low    |     med     |              low              |\n|       boosted-trees      |    both    |       low      |        high       |     low    |     high    |              low              |","metadata":{"id":"aKQ0aJKZnn5Y"}},{"cell_type":"markdown","source":"### Hyperparameter Tuning","metadata":{"id":"691LodTHnt7f"}},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\nfrom sklearn.exceptions import ConvergenceWarning\nimport warnings as warnings\n# KMeans couldn't produce obvious cluster, so will try some other models.\n# Going to focus on performance so will select SVM, multilater-perceptron, and boosted tree to compare.\n\"\"\"\n\"\"\"\ndef prep_tuning(self, clear_cache=False, ConvergenceWarning=ConvergenceWarning,\n                warnings=warnings, memory=joblib.Memory(AUGURY_CACHE, verbose=0)):\n  # Removing ConvergenceWarning from cells to keep cell results clean.\n  warnings.simplefilter('ignore', category=ConvergenceWarning)\n  if clear_cache:\n    memory.clear(warn=False)\ndef __call_fit(self, cv, X, y):\n  return cv.fit(X, y.values.ravel())\ndef feature_importance_report(self,model, X, np=np, plt=plt):\n  if not 'feature_importances_' in dir(model.best_estimator_):\n    return\n  feat_imp = model.best_estimator_.feature_importances_\n  indices = np.argsort(feat_imp)\n  plt.yticks(range(len(indices)), [X.columns[i] for i in indices])\n  plt.barh(range(len(indices)), feat_imp[indices], align='center')\n  plt.show()\ndef save_tuning_results(self, *, name, model):\n  models, = self._attr('models_')\n  if not name in models:\n    models[name] = {}\n  models[name]['model'] = model\n  self._attr(models_=models)\ndef hypertune_models(self, models, data=('X_train_', 'y_train_'), GridSearchCV=GridSearchCV, dt=dt,\n                     memory=joblib.Memory(AUGURY_CACHE, verbose=0),**kwargs):\n  assert len(models) <= 3, \"Max 3 Models allowed for Hyper Parameter Tuning\"\n  self.prep_tuning(**kwargs)\n  models_, = self._attr('models_')\n  for _m in models:\n    valid_keys = ['name', 'model', 'parameters']\n    _vals = self._hash(_m, *valid_keys)\n    for _i, _val in enumerate(_vals):\n      assert not _val is None, \"Missing {} key\".format(valid_keys[_i])\n    X, y = self._attr(*data)\n    name, model, parameters = _vals\n    GridSearchCV = memory.cache(GridSearchCV)\n    cv = GridSearchCV(model, parameters, cv=5)\n    print('{} ...GridSearchCV (this may take a while)\\n======='.format(name))\n    start_ = dt.now()\n    call_fit = memory.cache(self.__call_fit, ignore=['self'])\n    model_ = call_fit(cv, X, y)\n    print('Latency (hh:mm:ss.ms):{}'.format(dt.now() - start_))\n    print('Score (Mean Cross Validated): {}% | BEST PARAMS: {}\\n'.format(\n        round(model_.best_score_ * 100, 2), model_.best_params_))\n    self.feature_importance_report(model_, X)\n    self.save_tuning_results(name=name, model=model_)\nAugury.constructor()(prep_tuning,__call_fit,feature_importance_report, \n                     hypertune_models, save_tuning_results)\n\"\"\"\n\"\"\"","metadata":{"id":"0xFcsiOBVmeX","outputId":"684835ee-cda2-44b2-d4aa-d811789657bf","execution":{"iopub.status.busy":"2022-07-09T17:32:58.861429Z","iopub.execute_input":"2022-07-09T17:32:58.861972Z","iopub.status.idle":"2022-07-09T17:32:58.881602Z","shell.execute_reply.started":"2022-07-09T17:32:58.861928Z","shell.execute_reply":"2022-07-09T17:32:58.880521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\nfrom sklearn.neural_network import MLPClassifier, MLPRegressor\nfrom sklearn.ensemble import RandomForestClassifier, RandomForestRegressor\nfrom xgboost import XGBClassifier, XGBRegressor\n\"\"\"\nDefault Configs listed below\n\"\"\"\n## Note Regression has not been tested.\nclass_ = augur.label_type_ == 'categorical'\n##\nLogisticRegressionConfig = {\n  'name'      : 'LogisticRegression',\n  'model'     : LogisticRegression(),\n  'parameters': {\n    'C': [0.001, 0.01, 0.1, 1, 10, 100, 1000],\n    'class_weight' : [None, 'balanced']\n  }\n} if class_ else None\nSupportVectorMachinesConfig = {\n  'name'      : 'SupportVectorMachines',\n  'model'     : SVC(),\n  'parameters': {\n    'kernel': ['rbf', 'linear'],\n    'C'     : [0.1, 1, 10, 100]\n  },\n} if class_ else None\nMultilaterPerceptronConfig = {\n  'name'      : 'MultilaterPerceptron',\n  'model'     : MLPClassifier() if class_ else MLPRegressor(),\n  'parameters': {\n    'hidden_layer_sizes': [(10,), (50,), (100,)],\n    'activation'        : ['relu', 'tanh', 'logistic'],\n    'learning_rate'     : ['constant', 'invscaling', 'adaptive']\n  }\n}\nRandomForestConfig = {\n  'name'      : 'RandomForest',\n  'model'     : RandomForestClassifier() if class_ else RandomForestRegressor(),\n  'parameters': {\n    'n_estimators'  : [5, 50, 250, 500],\n    'max_depth'     : [1, 3, 5, 7, 9],\n    'learning_rate' : [0.01, 0.1, 1, 10, 100]\n  }\n}\nGradientBoostingConfig = {\n  'name'      : 'GradientBoosting',\n  'model'     : XGBClassifier() if class_ else XGBRegressor(),\n  'parameters': {\n    'n_estimators'  : [50, 250, 500],\n    'max_depth'     : [1, 3, 5, 7],\n    'learning_rate' : [0.01, 0.05, 0.1, 0.2],\n    'subsample'     : [0.8, 1],\n    'min_child_weight': [1, 10, 100]\n  }\n}\n\"\"\"\n\"\"\"\nMODEL_CONFIGS = [\n          LogisticRegressionConfig,\n          SupportVectorMachinesConfig,\n          #MultilaterPerceptronConfig, \n          #RandomForestConfig, \n          GradientBoostingConfig           \n]\naugur.hypertune_models(MODEL_CONFIGS)","metadata":{"id":"34LoQAFgYOBG","outputId":"026e45a7-3639-4832-aec2-a611648693ec","execution":{"iopub.status.busy":"2022-07-09T17:32:58.883328Z","iopub.execute_input":"2022-07-09T17:32:58.883949Z","iopub.status.idle":"2022-07-09T17:43:19.387317Z","shell.execute_reply.started":"2022-07-09T17:32:58.883905Z","shell.execute_reply":"2022-07-09T17:43:19.385743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fit","metadata":{"id":"BGPzQ_KdoN4D"}},{"cell_type":"code","source":"from sklearn.model_selection import learning_curve\n\"\"\"\n\"\"\"\ndef plot_learning_curve(self, models, plt=plt, dt=dt, learning_curve=learning_curve,\n                        memory=joblib.Memory(AUGURY_CACHE, verbose=0), np=np, \n                        data=('X_train_','y_train_')):\n  \"\"\"\n  Modified to display two Learning Curve graphs for side-by-side comparison,\n  from: https://scikit-learn.org/stable/auto_examples/model_selection/plot_learning_curve.html\n  \"\"\"\n  models_, X_train, y_train = self._attr('models_', *data)\n  _, axes = plt.subplots(1,3, figsize=(20,5))\n\n  start_ = dt.now()\n  idx = 0;\n  for model in models:\n    name = (lambda name, **_: name)(**model)\n    axes[idx].set_title(\"%s - Learning Curve\" % name)\n    axes[idx].set_xlabel(\"Training Examples\")\n    axes[idx].set_ylabel(\"Score\")\n    learning_curve = memory.cache(learning_curve)\n    train_sizes, train_scores, test_scores = learning_curve(\n        models_[name]['model'].best_estimator_,\n        X_train,\n        y_train,\n        cv=5,\n        train_sizes=np.linspace(0.1, 1.0, 5)\n    )\n    train_scores_mean = np.mean(train_scores, axis=1)\n    train_scores_std = np.std(train_scores, axis=1)\n    test_scores_mean = np.mean(test_scores, axis=1)\n    test_scores_std = np.std(test_scores, axis=1)\n\n    # Plot learning curve\n    axes[idx].grid()\n    axes[idx].fill_between(\n        train_sizes,\n        train_scores_mean - train_scores_std,\n        train_scores_mean + train_scores_std,\n        alpha=0.1,\n        color='r',\n    )\n    axes[idx].fill_between(\n        train_sizes,\n        test_scores_mean - test_scores_std,\n        test_scores_mean + test_scores_std,\n        alpha=0.1,\n        color='g',\n    )\n    axes[idx].plot(\n        train_sizes, train_scores_mean, \"o-\", color='r', label='Training Score'\n    )\n    axes[idx].plot(\n        train_sizes, test_scores_mean, 'o-', color='g', label='Cross-Validation Score'\n    )\n    axes[idx].legend(loc=\"best\")\n    idx += 1\n  plt.show()\n  print(\"Latency (hh:mm:ss.ms):{}\".format(dt.now() - start_))\nAugury.constructor()(plot_learning_curve)\naugur.plot_learning_curve(MODEL_CONFIGS)\n\"\"\"\n\"\"\"\n# View Model Performance","metadata":{"id":"YH5AXd9b4PAl","outputId":"631bf96f-1f30-4b93-e246-645daab53347","execution":{"iopub.status.busy":"2022-07-09T17:43:19.388873Z","iopub.execute_input":"2022-07-09T17:43:19.389190Z","iopub.status.idle":"2022-07-09T17:44:11.881483Z","shell.execute_reply.started":"2022-07-09T17:43:19.389161Z","shell.execute_reply":"2022-07-09T17:44:11.880734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Evaluation","metadata":{"id":"9Cyz-qU7oPjJ"}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, precision_score, recall_score\n\"\"\"\n\"\"\"\ndef evaluate_model(self, *, name, model, features, labels, dt=dt, average='binary',\n                   accuracy_score=accuracy_score, precision_score=precision_score,\n                   recall_score=recall_score, **kwargs):\n    start_ = dt.now()\n    pred = model.predict(features)\n    end_ = dt.now()\n    accuracy = accuracy_score(labels, pred)\n    precision = precision_score(labels, pred, average=average)\n    recall = recall_score(labels, pred, average=average)\n    self.set_performance(model=name, accuracy=accuracy, precision=precision,\n                           recall=recall, latency=end_-start_, **kwargs)\n    \ndef evaluate_models(self, model_configs, memory=joblib.Memory(AUGURY_CACHE, verbose=0), \n                    **kwargs):\n  models, X_train, y_train = self._attr('models_', 'X_train_', 'y_train_')\n  for model in model_configs:\n    name = (lambda name, **_: name)(**model)\n    evaluate_model = memory.cache(self.evaluate_model, ignore=['self'])\n    self.evaluate_model(name=name, model=models[name]['model'].best_estimator_, \n                        features=X_train, labels=y_train, **kwargs)\nAugury.constructor()(evaluate_models, evaluate_model)\naugur.evaluate_models(MODEL_CONFIGS, output_string=True)\n\"\"\"\n\"\"\"\n# Evaluate performance of models against Valdation Data Set","metadata":{"id":"GoRw4CSu-oHh","outputId":"6b108c41-0edc-4944-c033-ce214626e918","execution":{"iopub.status.busy":"2022-07-09T17:44:11.882708Z","iopub.execute_input":"2022-07-09T17:44:11.883192Z","iopub.status.idle":"2022-07-09T17:44:13.501185Z","shell.execute_reply.started":"2022-07-09T17:44:11.883162Z","shell.execute_reply":"2022-07-09T17:44:13.500382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"EVALUATION_CRITERIA = {\n    'criteria': \"accuracy\" # one of (accuracy|precision|recall)\n}\n\"\"\"\n\"\"\"\ndef set_evalutation_criteria(self, *, criteria):\n  _valid = ['accuracy', 'precision', 'recall']\n  assert criteria in _valid, \"Invalid criteria, expecting one of {}.\".format(_valid)\n  self._attr(evaluation_criteria_=criteria)\nAugury.constructor()(set_evalutation_criteria)\naugur.set_evalutation_criteria(**EVALUATION_CRITERIA)\n\"\"\"\n\"\"\"\naugur.evaluation_criteria_","metadata":{"id":"vuObjZWZKEWw","outputId":"bed4d34e-1698-42a6-d7b9-453a68e5437a","execution":{"iopub.status.busy":"2022-07-09T17:44:13.502414Z","iopub.execute_input":"2022-07-09T17:44:13.503498Z","iopub.status.idle":"2022-07-09T17:44:13.512563Z","shell.execute_reply.started":"2022-07-09T17:44:13.503455Z","shell.execute_reply":"2022-07-09T17:44:13.511733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Tune\nWill compare against compressed dataset.","metadata":{"id":"y_C1H2MB2q9G"}},{"cell_type":"code","source":"\"\"\"\n#TODO\ndef tune(self):\n  evaluation_criteria = getattr(self, 'evaluation_criteria')\n  model_name = getattr(self, \"best_%s\" % evaluation_criteria)[0]\n  # Todo - need to refactor some of the above to be a bit less encompassing\nfor _func in [tune]:\n  constructor_(Augury)(_func)\n\n\"\"\"","metadata":{"id":"hevv93woLpFd","outputId":"4a5fcd76-dd03-4e24-a292-a5934e6e8850","execution":{"iopub.status.busy":"2022-07-09T17:44:13.513916Z","iopub.execute_input":"2022-07-09T17:44:13.514453Z","iopub.status.idle":"2022-07-09T17:44:13.527035Z","shell.execute_reply.started":"2022-07-09T17:44:13.514423Z","shell.execute_reply":"2022-07-09T17:44:13.526199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test\nRead in File\nApply same transformation as training, and predict with best model.","metadata":{"id":"MrMeU5Yu-uWZ"}},{"cell_type":"code","source":"\"\"\"\n\"\"\"\naugur.set_data_from_path(data_test_=TEST_PATH)\n\"\"\"\n\"\"\"\n# View Test data\naugur.data_test_","metadata":{"id":"WvWrT04Xf7y4","outputId":"40867147-1b9b-46cd-dc82-59efd7b05233","execution":{"iopub.status.busy":"2022-07-09T17:44:13.528243Z","iopub.execute_input":"2022-07-09T17:44:13.528753Z","iopub.status.idle":"2022-07-09T17:44:13.575370Z","shell.execute_reply.started":"2022-07-09T17:44:13.528724Z","shell.execute_reply":"2022-07-09T17:44:13.574598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\"\"\"\nTEST_DATA_ATTRS = ('data_test_clean_', 'data_test_') # Hardcoded for now\ndef prepare_predictions(self, pd=pd, id_column=None):\n    data_test, = self._attr('data_test_')\n    assert not data_test is None\n    id_column = data_test.columns[0] if id_column is None else id_column\n    self._attr(predictions_=pd.DataFrame(data_test[id_column]))\n    return self\nAugury.constructor()(prepare_predictions)\naugur.prepare_predictions().process_nulls(\n    data=TEST_DATA_ATTRS).process_features(data=TEST_DATA_ATTRS).process_continuous_features(\n        data=TEST_DATA_ATTRS).clip_continuous(data=TEST_DATA_ATTRS)\n\"\"\"\n\"\"\"\n# View cleaned test data\naugur.data_test_clean_","metadata":{"id":"BA9auY3Recg_","outputId":"350a308f-58e1-49d8-e29b-96399c6713c0","execution":{"iopub.status.busy":"2022-07-09T17:44:13.576453Z","iopub.execute_input":"2022-07-09T17:44:13.576916Z","iopub.status.idle":"2022-07-09T17:44:14.247816Z","shell.execute_reply.started":"2022-07-09T17:44:13.576890Z","shell.execute_reply":"2022-07-09T17:44:14.246816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\"\"\"\ndef process_boxcox(self):\n  boxcox_pipeline, data_test_clean= self._attr('boxcox_pipeline_', 'data_test_clean_')\n  for rule in boxcox_pipeline:\n    feature, lmbda = rule\n    data_test_clean[\"%s_clean_tr\" % feature] = self.__call_boxcox(\n        lmbda=lmbda, data=data_test_clean, column=\"%s_clean\" % feature)\n  return self\nAugury.constructor()(process_boxcox)\naugur.process_boxcox().numeric_transformation(data='data_test_clean_')\n\"\"\"\n\"\"\"\n# View transformed data\naugur.data_test_clean_","metadata":{"id":"wjV2MF2papGb","outputId":"22453e46-e602-45ca-e2e1-9180e4319ff3","execution":{"iopub.status.busy":"2022-07-09T17:44:14.249229Z","iopub.execute_input":"2022-07-09T17:44:14.249752Z","iopub.status.idle":"2022-07-09T17:44:14.316279Z","shell.execute_reply.started":"2022-07-09T17:44:14.249720Z","shell.execute_reply":"2022-07-09T17:44:14.315539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\"\"\"\naugur.set_data_final(test=True)\n\"\"\"\n\"\"\"\n# View test data to be predicted against\naugur.data_test_final_","metadata":{"id":"AGU8zOg1dfzy","outputId":"e2fa32a6-43ef-47ae-d920-b49c48db5fe5","execution":{"iopub.status.busy":"2022-07-09T17:44:14.317497Z","iopub.execute_input":"2022-07-09T17:44:14.318059Z","iopub.status.idle":"2022-07-09T17:44:14.356663Z","shell.execute_reply.started":"2022-07-09T17:44:14.318029Z","shell.execute_reply":"2022-07-09T17:44:14.355494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\"\"\"\ndef set_scaled_test(self):\n  data_test_final, scaler= self._attr('data_test_final_', 'scaler_')\n  self._attr(data_test_scaled_=pd.DataFrame(scaler.transform(data_test_final),\n                                           index=data_test_final.index,\n                                           columns=data_test_final.columns))  \nAugury.constructor()(set_scaled_test)\naugur.set_scaled_test()\n\"\"\"\n\"\"\"\naugur.data_test_scaled_","metadata":{"id":"mzXAy99XdWWX","outputId":"e5cc801f-a48f-463f-cc5d-f882da89c10f","execution":{"iopub.status.busy":"2022-07-09T17:44:14.357821Z","iopub.execute_input":"2022-07-09T17:44:14.358571Z","iopub.status.idle":"2022-07-09T17:44:14.387536Z","shell.execute_reply.started":"2022-07-09T17:44:14.358541Z","shell.execute_reply":"2022-07-09T17:44:14.386748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\"\"\"\ndef predict(self, pd=pd, dt=dt):\n  evaluation_criteria, models, predictions, label, data_test_scaled = \\\n    self._attr('evaluation_criteria_', 'models_', 'predictions_', 'label_', 'data_test_scaled_')\n  (model_name, _), = self._attr(\"best_%s_\" % evaluation_criteria) # Nested Tuple\n  model = models[model_name]['model']\n  start_ = dt.now()\n  _predictions = model.predict(data_test_scaled)\n  print(\"Latency (hh:mm:ss.ms):{}\".format(dt.now() - start_))\n  predictions[label] = _predictions\nAugury.constructor()(predict)\naugur.predict()\n# For predictions to work from Kaggle\naugur.predictions_[augur.label_] = augur.predictions_[augur.label_].astype(bool)\n\"\"\"\n\"\"\"\naugur.predictions_","metadata":{"id":"P2yDGjWRLa1j","outputId":"fe4691bb-8332-4e98-907a-29b57e43bf74","execution":{"iopub.status.busy":"2022-07-09T17:44:14.388627Z","iopub.execute_input":"2022-07-09T17:44:14.389261Z","iopub.status.idle":"2022-07-09T17:44:14.421839Z","shell.execute_reply.started":"2022-07-09T17:44:14.389227Z","shell.execute_reply":"2022-07-09T17:44:14.420728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"\n\"\"\"\ndef prepare_submission(self, submission_path='./submission.csv'):\n  predictions, = self._attr('predictions_')\n  predictions.to_csv(submission_path, index=False)\n  self._attr(submission_path_=submission_path)\n  print(\"Predictions exported to:{}\".format(submission_path))\nAugury.constructor()(prepare_submission)\naugur.prepare_submission()\n\"\"\"\n\"\"\"\n# Predictions csv created.","metadata":{"id":"QPSz8zmulgvW","outputId":"adfbdb0f-d448-452d-b60c-7f4a55ae427c","execution":{"iopub.status.busy":"2022-07-09T17:44:14.423283Z","iopub.execute_input":"2022-07-09T17:44:14.424343Z","iopub.status.idle":"2022-07-09T17:44:14.442147Z","shell.execute_reply.started":"2022-07-09T17:44:14.424303Z","shell.execute_reply":"2022-07-09T17:44:14.441093Z"},"trusted":true},"execution_count":null,"outputs":[]}]}