{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Generic CNN Training Pipeline\nPresented here is a pipeline for doing a single round of hyperband hyperparameter optimization on a single standard CNN architecture (with problem-specific classification head). While the results obviously won't be as impressive as those achieved by a notebook that trains and ensembles _multiple_ CNN architectures, the hope is to build some code that can be used to arrive at good hyperparameter choices with which to seed such a notebook.\n\nIn practice, since hyperparameters aren't required to be optimized in the notebook for code competitions (we can pretend we arrived at them by lucky guessing or magic), such a pipeline would probably be best run offline on some personal compute resources, since the 9 hour time constraint limits the size of models we can optimize over and the types of data pipelines we can support (using `albumentations`, for example, in this setting would bottleneck our training step on CPU compute and keep us from doing an exhaustive search).\n\nIn contrast to a lot of notebooks in this challenge, which tend to leverage PyTorch, I plan on leveraging TensorFlow in this notebook, just because I love Keras and much of the surrounding ecosystem, and because I feel like trying something different. As we'll see, however, the cost of whatever code-cleanliness Keras provides is a lack of robustness in being able to train multiple models in the same process. This will motivate us to launch training jobs as subprocesses, and only do the hyperparameter optimization/final prediction in this process.\n\nWhile this ultimately incurs a non-trivial amount code overhead, the benefit is that if we're clever about how we solve things, much of what we build here will be useful for similar challenges down the line, empowering us to focus on the TF/DL side of things.\n\n## Getting started\nStart with our imports and define some global variables","metadata":{}},{"cell_type":"code","source":"import copy\nimport glob\nimport inspect\nimport json\nimport os\nimport random\nimport re\nimport subprocess\nimport time\nimport typing\nfrom functools import partial, wraps\n\nimport attr\nimport kerastuner as kt\nimport pandas as pd\nimport tensorflow as tf\nfrom sklearn.model_selection import KFold\n\n\n# some common paths\nDATA_DIR = \"/kaggle/input/cassava-leaf-disease-classification\"\nWEIGHTS_DIR = \"/kaggle/input/pretrained-weights\"\nOUTPUT_DIR = \"/kaggle/working\"\n\n# some useful constants\nNUM_SAMPLES = 21397\nIMAGE_SHAPE = (512, 512, 3)\nMEAN = tf.constant([109.9674 , 129.97475,  78.19961], dtype=tf.float32)\nSTD = tf.constant([59.77377 , 60.50442 , 57.367416], dtype=tf.float32)\nTEST = False\nSUBMIT = True\nSTART_TIME = time.time() # TODO: use this for deciding whether to run more rounds","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Utilities\nI'm going to define this `RunParams` object in order to standardize things like naming conventions and model instantiation, especially since the params I want to use are still in flux and this helps with updating and keeping tabs on them. It's gross overkill, but I've just recently discovered the `attr` library and feel like playing with it. Hopefully it will make hyperparameter searching a bit simpler.","metadata":{}},{"cell_type":"code","source":"\ndef data_path(*args):\n    \"\"\"\n    Quick utility function so I don't have to keep\n    typing paths or os.path.join over and over\n    \"\"\"\n    return os.path.join(DATA_DIR, *args)\n\n\ndef output_path(*args):\n    \"\"\"\n    Same thing but for saving stuff\n    \"\"\"\n    return os.path.join(OUTPUT_DIR, *args)\n\n\ndef box_print(msg: str, num_spaces: int):\n    \"\"\"\n    Quick utility for pretty printing messages\n    in a hacky little dedicated box,\n    -------------------\n    |    like this    |\n    -------------------\n    \"\"\"\n    # msg + 2*(padding + vertical bars)\n    num_dashes = len(msg) + 2*(num_spaces+1)\n    print(\"\\n\" + \"-\"*num_dashes)\n    print(\"|\" + \" \"*num_spaces + msg + \" \"*num_spaces + \"|\")\n    print(\"-\"*num_dashes)\n\n\n# TODO: How do we replace the RunParams entirely\n# with a kerastuner Hyperparameters instance that\n# just treats all non-searchable parameters as\n# `Fixed` type hyperparameters\ndef float_attrib(n_digits: int, default: typing.Optional[float] = None):\n    \"\"\"\n    dumb wrapper for string formmating float params\n    \"\"\"\n    metadata = {\"formatter\": \"{{:0.{}f}}\".format(n_digits)}\n    kwargs = {\"metadata\": metadata}\n    if default is not None:\n        kwargs[\"default\"] = default\n    return attr.ib(**kwargs)\n\n\nclass ParamConfig:\n    @classmethod\n    def from_string(cls, string):\n        props = dict([i.split(\"=\") for i in string.split(\"-\")])\n\n        for attr in cls.__attrs_attrs__:\n            try:\n                type_name = attr.type._name\n            except AttributeError:\n                type_name = None\n\n            if type_name == \"List\":\n                type_map = attr.type.__args__[0]\n                props[attr.name] = list(\n                    map(type_map, props[attr.name].split(\",\"))\n                )\n            else:\n                props[attr.name] = attr.type(props[attr.name])\n        return cls(**props)\n\n    @classmethod\n    def from_json(cls, json_path):\n        with open(json_path, \"r\") as f:\n            return cls(**json.load(f))\n\n    def to_json(self, json_path):\n        with open(json_path, \"w\") as f:\n            json.dump(self.__dict__, f)\n\n    def __str__(self):\n        string = \"\"\n        for attr in self.__attrs_attrs__:\n            value = self.__dict__[attr.name]\n            try:\n                formatter = attr.metadata[\"formatter\"]\n            except KeyError:\n                try:\n                    if attr.type._name == \"List\":\n                        formatter = lambda x: \",\".join(\n                            list(map(str, x))\n                        )\n                except AttributeError:\n                    formatter = str\n\n            if not isinstance(formatter, typing.Callable):\n                format_string = formatter\n                formatter = lambda x: format_string.format(x)\n            string += attr.name + \"=\" + formatter(value) + \"-\"\n        return string[:-1]\n\n    def copy(self):\n        return copy.deepcopy(self)\n\n\n@attr.s(auto_attribs=True)\nclass RunParams(ParamConfig):\n    \"\"\"\n    Config describing a particular model to\n    train and the hyperparameters (optimized\n    or otherwise) used to fit it to data.\n    Use this to keep from exposing the same\n    args over and over to the functions we\n    build on top of our fitting function\n    \"\"\"\n    backbone: str\n    hidden_dims: typing.List[int]\n    downscale: int\n    crop_factor: float = float_attrib(4)\n    batch_size: int\n    learning_rate: float = float_attrib(5)\n    decay_steps: int\n    epochs: int\n    patience: int = 3\n    label_smoothing: float = float_attrib(3, default=0)\n    contrast: float = float_attrib(3, default=0.2)\n    rotation: float = float_attrib(3, default=0.2)\n    zoom: float = float_attrib(3, default=0.2)\n    blue_noise: float = float_attrib(3, default=0.2)\n\n    @property\n    def input_shape(self):\n        return (\n            tuple([i//self.downscale for i in IMAGE_SHAPE[:2]]) + \n            (IMAGE_SHAPE[2],)\n        )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset Utilities\nSince the data is made available as TFRecords, we'll leverage TF's dataset class to load, parse, and possibly downscale the images. Unfortunately, TF's JPEG parser can't be run in batch, and I'm suspicious that this data loading step ultimately bottlenecks the pipeline (though I really ought to leverage TensorBoard's profiling tool to confirm this). Note that besides the JPEG downsizing, all of the augmentation will be handled by Keras' preprocessing layers, which help move more compute onto the GPU.","metadata":{}},{"cell_type":"code","source":"def parse_fn(\n    example_proto,\n    downscale: int = 4,\n    return_filenames: bool = False,\n    label: bool = True,\n    one_hot: bool = True\n) -> typing.Tuple[\n        typing.Union[tf.Tensor, typing.Dict[str, tf.Tensor]],\n        tf.Tensor\n]:\n    \"\"\"\n    parses a tfrecord with the jpeg image data,\n    decodes the jpeg, and casts it to fp32. In\n    an ideal world, this would be done on a batch\n    of protos using `tf.io.parse_example`, but the\n    jpeg decode can't be done in batch.\n    TODO: see if we can leverage Dali for data loading.\n\n    If `label` is `True`, returns a tuple of `(input, target)`\n    pairs, where `input` is either a `tf.Tensor` corresponding\n    to the image, or a `dict` containing both this tensor\n    and the corresponding image filename if `return_filenames`\n    is `True`.\n\n    :param example_proto: Serialized `tf.train.Example`\n        protocol buffer containing image data\n    :param downscale: Factor by which to downscale the\n        image jpeg at decoding time. Faster than doing\n        it during preprocessing if you're comfortable\n        using the available power of 2 factors\n    :param return_filenames: Whether to include the\n        filename of the corresponding image in the\n        model input, returning a `dict` as the\n        first return element instead of a `tf.Tensor`\n    :param label: Whether to return the class label\n        for the given `Example` (if it has one).\n    :param one_hot: If returning the class label,\n        whether to return it as an `int` or convert\n        it to a \"one hot\" binary vector representation.\n    \"\"\"\n    # parse the protobuf byte string into tf Tensors\n    # using the feature spec\n    feature_spec = {\n        \"image\": tf.io.FixedLenFeature([], tf.string),\n        \"image_name\": tf.io.FixedLenFeature([], tf.string)\n    }\n    if label:\n        feature_spec[\"target\"] = tf.io.FixedLenFeature([], tf.int64)\n    tensors = tf.io.parse_single_example(example_proto, feature_spec)\n\n    # convert the jpeg byte string to a 3D Tensor\n    # normalize it with the training set stats\n    # if we're returning filenames, then make the\n    # whole input a dict with the image and filename\n    x = tf.io.decode_jpeg(tensors[\"image\"], ratio=downscale)\n    x = (tf.cast(x, tf.float32) - MEAN) / STD\n    if return_filenames:\n        x = {\"raw_input\": x, \"image_name\": tensors[\"image_name\"]}\n\n    # return labels if need be, possibly one-hot encoding\n    if label:\n        y = tensors[\"target\"]\n        if one_hot:\n            y = tf.one_hot(y, 5)\n        return x, y\n    return x\n\n\ndef make_dataset(\n    fnames: typing.List[str],\n    batch_size: int = 32,\n    shuffle: bool = True,\n    downscale: int = 4,\n    return_filenames: bool = False,\n    label: bool = True,\n    one_hot: bool=True\n) -> tf.data.Dataset:\n    \"\"\"\n    Builds a TF `Dataset` that iterates through the\n    provided tfrecord `fnames`.\n\n    :param fnames: Filenames to iterate through\n    :param batch_size: Number of samples to return at\n        each iteration\n    :param shuffle: Whether to return the data in a\n        random order\n    :param downscale: Factor by which to downscale the\n        image data at decode time\n    :param return_filenames: Whether to include the\n        corresponding filename as a field with each\n        image. See the documentation for `parse_fn`\n    :param label: Whether to include class labels\n        with the images (if they're available)\n    :param one_hot: Whether to return class labels\n        as integers or encoded into binary \"one-hot\"\n        vectors\n    \"\"\"\n    dataset = tf.data.TFRecordDataset(fnames)\n    if shuffle:\n        dataset = dataset.shuffle(\n            buffer_size=10000, reshuffle_each_iteration=True\n        )\n\n    _parse_fn = partial(\n        parse_fn,\n        downscale=downscale,\n        return_filenames=return_filenames,\n        label=label,\n        one_hot=one_hot\n    )\n    return dataset.map(\n        partial(_parse_fn),\n        num_parallel_calls=tf.data.experimental.AUTOTUNE\n    ).batch(batch_size) # .prefetch(tf.data.experimental.AUTOTUNE)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Custom Preprocessing\nAs I note in [this notebook](https://www.kaggle.com/alecgunny/cassava-visualization?scriptVersionId=51068011), most of the color differences we're interested in happen along axes orthogonal to the blue channel of the image (i.e. green, yellow, and brown all differ primarily in the relative amount of red/green content). To take advantage of this, we'll make a custom random augmentation that adds noise to _just_ the blue channel, since we want to make sure the network learns that this content is largely irrelevant for its decision making. I'm still not convinced this is the best idea (and even if the concept _is_ good there may be better ways to exploit this invariance), but in the spirit of trying different things and showing the extensibility of Keras' APIs we'll start with this.\n\nWe'll do this with a subclassed Keras `PreprocessingLayer`, drawing a standard deviation for the noise from a uniform distribution, then multiplying it by random gaussian noise. Due to the way TensorFlow handles (or doesn't handle) tensor assignment, the actual implementation generates a gaussian noise tensor for all channels, then gets rid of the red and green contributions with a binary mask generated at build time.","metadata":{}},{"cell_type":"code","source":"from tensorflow.python.keras.utils.control_flow_util import smart_cond\n\n\n# this lets us leverage this layer with tf SavedModel\n@tf.keras.utils.register_keras_serializable(name=\"RandomBlueNoise\")\nclass RandomBlueNoise(\n        tf.keras.layers.experimental.preprocessing.PreprocessingLayer\n):\n    \"\"\"\n    Layer for adding Gaussian noise with a randomly\n    selected standard deviation to the blue channel\n    of an RGB image.\n\n    :params max_std: The maximum standard deviation\n        to be randomly selected for the noise\n    \"\"\"\n    def __init__(\n        self,\n        max_std: float,\n        **kwargs\n    ):\n        kwargs[\"trainable\"] = False\n        super(RandomBlueNoise, self).__init__(**kwargs)\n        self.max_std = max_std\n        self.input_spec = tf.keras.layers.InputSpec(ndim=4)\n\n    def build(self, input_shapes):\n        self.mask = self.add_weight(\n            name=\"mask\",\n            shape=input_shapes[1:],\n            initializer=\"zeros\",\n            trainable=False,\n            dtype=tf.float32\n        )\n        self.mask[:, :, -1].assign(\n            tf.ones(input_shapes[1:-1])\n        )\n\n    def call(self, inputs, training=False):\n        if training is None:\n            training = tf.keras.backend.learning_phase()\n        shape = tf.shape(inputs)\n\n        def add_blue_noise():\n            outputs = inputs\n            stds = tf.random.uniform(\n                shape=[shape[0]], maxval=self.max_std\n            )\n            noise = tf.random.normal(shape)\n            noise = tf.transpose(noise, (1, 2, 3, 0))*stds\n            noise = tf.transpose(noise, (3, 0, 1, 2))\n            noise = noise*self.mask\n            outputs = outputs + noise\n            return outputs\n\n        output = smart_cond(\n            training, add_blue_noise, lambda: inputs\n        )\n        output.set_shape(inputs.shape)\n        return output\n\n    def compute_output_shape(self, input_shape):\n        return input_shape\n\n    def get_config(self):\n        config = {\"max_std\": self.max_std}\n        base_config = super(RandomBlueNoise, self).get_config()\n        return dict(list(base_config.items()) + list(config.items()))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The Model\nNow that we have our custom tools ready, we'll define a generic CNN by doing some random preprocessing augmentations (including the custom one we just created), doing feature extraction from the preprocessed images via some common pre-trained convolutional backbone network, and finally using a dense classifier to map those features to a logistic output.","metadata":{}},{"cell_type":"code","source":"def make_preprocessor(\n    height: int,\n    width: int,\n    contrast: float,\n    rotation: float,\n    zoom: float,\n    blue_noise: float\n):\n    return tf.keras.Sequential([\n        RandomBlueNoise(blue_noise),\n        tf.keras.layers.experimental.preprocessing.RandomContrast(\n            contrast\n        ),\n        tf.keras.layers.experimental.preprocessing.RandomRotation(\n            (-rotation, rotation)\n        ),\n        tf.keras.layers.experimental.preprocessing.RandomZoom(\n            (-zoom, zoom), fill_mode=\"constant\"\n        ),\n        tf.keras.layers.experimental.preprocessing.RandomCrop(\n            width=width, height=height\n        ),\n        tf.keras.layers.experimental.preprocessing.RandomFlip()\n    ], name=\"preprocessor\")\n\n\ndef make_feature_extractor(backbone: str, height: int, width: int):\n    # check to make sure we have weights\n    available_weights = tf.io.gfile.listdir(WEIGHTS_DIR)\n    if backbone not in [i.split(\".\")[0] for i in available_weights]:\n        raise ValueError(f\"No pretrained weights for backbone {backbone}\")\n    else:\n        weights = os.path.join(WEIGHTS_DIR, f\"{backbone}.h5\")\n\n    # get prebuilt class\n    backbone_classname = \"\".join([i.title() for i in backbone.split(\"_\")])\n    try:\n        backbone = getattr(tf.keras.applications, backbone_classname)\n    except AttributeError:\n        raise ValueError(f\"Unknown backbone {backbone}\")\n\n    # instantiate and call on output of preprocessor model\n    return backbone(\n        input_shape=(height, width, 3),\n        include_top=False,\n        pooling=\"avg\",\n        weights=weights\n    )\n\n\ndef make_classifier(hidden_dims: typing.List[int]):\n    # use this to get how many output classes\n    # we need rather than hard coding it for\n    # the sake of aesthetics\n    label_map_path = data_path(\"label_num_to_disease_map.json\")\n    with open(label_map_path, \"r\") as f:\n        num_classes = len(json.load(f))\n\n    layers = [\n        tf.keras.layers.Dense(dim, activation=\"relu\") for dim  in hidden_dims\n    ]\n    layers.append(tf.keras.layers.Dense(num_classes, activation=\"softmax\"))\n    return tf.keras.Sequential(layers, name=\"classifier\")\n\n\ndef generic_cnn(\n    backbone: str,\n    input_shape: typing.Tuple[int, int, int] = IMAGE_SHAPE,\n    classifier_hidden_dims: typing.List[int] = [1024],\n    crop_factor: float = 0.8,\n    contrast: float = 0.2,\n    rotation: float = 0.2,\n    zoom: float = 0.2,\n    blue_noise: float = 0.5\n) -> tf.keras.Model:\n    \"\"\"\n    Build a Keras Model that performs random preprocessing,\n    builds feature maps from a pretrained convolutional backbone,\n    and uses dense layers to map these to class probabilities.\n\n    :param backbone: name of pre-trained CNN feature extractor\n        to use. Options right now are `'res_net_50_v2'` and\n        `'mobile_net_v2'`.\n    :param input_shape: tuple indicating shape of an image\n        in the batch (this means _after_ any resizing done\n        during data loading), with the channel dimension\n        last. Defaults to the full image shape `(512, 512, 3)`.\n    :param classifier_hidden_dims: list of sizes for the\n        hidden layers in the classifier. Note that there will\n        always be an additional output layer with dimension 5.\n    :param crop_factor: fraction of image spatial dimensions\n        to crop, rounded to the nearest integer number of pixels.\n    :param contrast: amount of random contrast augmentation to\n        apply\n    :param rotation: range of random rotations to select from,\n        as a fraction of 360 degrees. Range will be\n        `(-rotation, rotation)`.\n    :param zoom: range of random zoom factors to select from\n    :param blue_noise: range of randomly selected standard\n        deviation for gaussian noise to be added to blue\n        channel of each image in batch\n    \"\"\"\n    # input shapes to feature extractor after cropping\n    width = int(crop_factor*input_shape[1])\n    height = int(crop_factor*input_shape[0])\n\n    # start by building a preprocessor model\n    # that does a bunch of random augmentations,\n    # including injecting blue noise\n    preprocessor = make_preprocessor(\n        height,\n        width,\n        contrast,\n        rotation,\n        zoom,\n        blue_noise\n    )\n\n    # next build the convolutional part of the\n    # network, which maps this preprocessed\n    # input to some learned feature space\n    # used pre-trained weights to avoid having\n    # to do the onerous work of learning things\n    # like edges, shapes, etc.\n    feature_extractor = make_feature_extractor(\n        backbone, height, width\n    )\n\n    # finally build a dense classifier which\n    # leverages ReLU activations to map from these\n    # learned features to a logistic output\n    classifier = make_classifier(classifier_hidden_dims)\n\n    # finally, instantiate an input tensor and\n    # then call each section of the network on\n    # it as if it's just another layer\n    raw_input = tf.keras.Input(name=\"input\", shape=input_shape)\n    x = preprocessor(raw_input)\n    x = feature_extractor(x)\n    x = classifier(x)\n    return tf.keras.Model(inputs=raw_input, outputs=x)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prediction\nBefore I define my training function, I'll define a quick function for producing predictions on dataset so that we can do validation at the end.","metadata":{}},{"cell_type":"code","source":"def predict_on_dataset(\n    model: tf.keras.Model,\n    dataset: tf.data.Dataset,\n    training: bool = True\n) -> pd.DataFrame:\n    \"\"\"\n    Produce a dataframe of predictions on a\n    dataset using the given model, optionally\n    including a `\"label\"` column for the label\n    if it's in the given dataset\n\n    :param model: Keras model to use for predictions\n    :param dataset: Dataset to produce predictions on\n    :param training: Flag to pass to `model.__call__`\n        to indicate whether to perform augmentations\n        or not\n    \"\"\"\n    predictions = []\n    y, image_id = None, None\n    for X in dataset:\n        if not isinstance(X, (dict, tf.Tensor)):\n            # presumably a tuple, indicating that\n            # we have a y to look for\n            X, y = X\n            y = y.numpy()\n\n            # check if we have a one-hot, in which case\n            # take the argmax for the integer label\n            if len(y.shape) > 1:\n                y = y.argmax(axis=1)\n\n        # if we have the image id included in the\n        # input, then take it to make X a tf Tensor\n        if isinstance(X, dict):\n            image_id = X[\"image_name\"].numpy()\n            X = X[\"raw_input\"]\n        assert isinstance(X, tf.Tensor)\n\n        # make our predictions and shove them\n        # into a dataframe\n        y_hat = model(X, training=training)\n        df = pd.DataFrame(y_hat.numpy())\n\n        # include labels and metadata if we have them\n        if y is not None:\n            df[\"label\"] = y.flatten()\n        if image_id is not None:\n            df[\"image_id\"] = image_id\n            df[\"image_id\"] = df[\"image_id\"].str.decode(\"utf-8\")\n            df = df.set_index(\"image_id\")\n\n        # append and continue\n        predictions.append(df)\n\n    # return our concatenated predictions\n    return pd.concat(predictions, axis=0)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Custom Callback\nWhile we're hyperparameter searching, we're bound to stumble upon combinations that are going nowhere. Rather than waiting the full `patience` number of steps to see if they course correct, we'll build a custom `EarlyStopping` callback that uses a baseline value to decide if training is even worth seeing through. This is super easy by just subclassing the existing callback.","metadata":{}},{"cell_type":"code","source":"class HardBaselineEarlyStopping(tf.keras.callbacks.EarlyStopping):\n    def on_epoch_end(self, epoch, logs=None):\n        current = self.get_monitor_value(logs)\n        if current < self.baseline:\n            self.model.stop_training = True\n        else:\n            super().on_epoch_end(epoch, logs)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training Pipeline\nNow we can stitch all these elements together into a full training run for a given dataset and mdoel. But first, there's a bit of an annoying catch...\n\n### A Quick Detour\nIn a moment, we'll define a function that will build and train a model on some subset of the dataset, then validate on some other subset. Given the size of the dataset, holding out a dedicated validation set takes away an already precious resource. For that reason, we'll do a full cross validation on these training runs in order to use as much data as possible (at the expense of a lot of compute).\n\nMoreover, I have no idea if my chose hyperparameter (HP) settings are any good, so it would be great to do an HP search using this cross validation score as the objective. However, TensorFlow has a bit of a memory leak that keeps us from being able to perform even the cross validation without running out memory, let alone _many_ cross validations. (I still don't know exactly where the memory leak is coming from. I used to suspect it was coming from poor garbage collection on the dataset iterators, but a lot of experimentation has led me to believe it's actually coming from the model itself).\n\nThe good news is that everything I've done so far is just defining functions: I haven't actually gone and allocated any GPU memory for TensorFlow yet. So instead of doing my cross validation/HP search in this kernel, I'll launch each training as a subprocess with its _own_ kernel. This way, each job uses all the memory its little heart desires, and then gives it back. We'll pay for this with some launch overhead for each job, but if the jobs themselves are long enough this should be a relatively small price to pay.\n\nHow we'll do this is with a bit of Jupyter Notebook hacking. Basically, I'll create an `if __name__ == \"__main__\":` section below, and tag it with a `# last cell` flag at the start. I'll then use Jupyter Notebook's cell content cache to iterate through all of these cells, stopping at the one below, taking their text, and dumping it into a python file. I can then _call_ that python file in a subprocess, and parse its output to get the validation accuracy measurement. I'll even define a fun utility function below that will take our training function and wrap with a function that takes the same arguments, but calls the script as a subprocess. If you think this is overkill, I'll mention that I used to use the same functionality for the cross validation, instead of each training job separately, until I realized that a cross validation couldn't run to completion. Luckily, I was able to use this function to switch from one to the other without too much overhead.\n\nI wouldn't call this good coding practice exactly, but it solves our problem pretty neatly and has the added benefit of giving us a training script which should hopefully be pretty useful more generally. The broader lesson is probably that TensorFlow, while it has the superior API in my view and is well suited to hyperscale ML use cases (at least if you're using GCP), doesn't necessarily extend its behavior to simple but slightly different use cases very well.","metadata":{}},{"cell_type":"code","source":"@attr.s(auto_attribs=True)\nclass scriptify:\n    \"\"\"\n    Function decorator that will wrap normal\n    functions in their script-calling equivalent.\n    Uses a `wrap` kwarg to use real behavior\n    if being called in the subprocess, and a `callback`\n    function to return the appropriate results.\n\n    TODO: is there a way to do emulate the `wrap`\n    functionality by just using the PID?\n    TODO: can we create a similar function that\n    returns the argument parser as well?\n    \"\"\"\n    wrap: bool = True\n    callback: typing.Optional[typing.Callable] = None\n    def __call__(self, func):\n        if not self.wrap:\n            return func\n\n        @wraps(func)\n        def wrapper(*args, **kwargs):\n            cmd = [\"python\", \"main.py\"]\n\n            # params is useful in case we need\n            # defaults, but param_names gives\n            # us everything in a list (aka in order)\n            params = inspect.signature(func).parameters\n            param_names = func.__code__.co_varnames[\n                :func.__code__.co_argcount\n            ]\n            for idx, argname in enumerate(param_names):\n                # start by trying to get the value\n                # passed to the current arg\n                try:\n                    # try for args first\n                    value = args[idx]\n                except IndexError:\n                    try:\n                        # see if it was in a kwarg maybe\n                        value = kwargs[argname]\n                    except KeyError:\n                        # maybe it has a default?\n                        if params[argname].default is inspect._empty:\n                            # ok we must have screwed up\n                            raise ValueError(\n                                f\"No value provided for arg {arg}\"\n                            )\n                        value = params[argname].default\n\n                if (\n                    isinstance(value, bool) and not value\n                    or value is None\n                ):\n                    continue\n                cmd.append(\"--\" + argname.replace(\"_\", \"-\"))\n\n                if isinstance(value, list):\n                    cmd.extend(map(str, value))\n                elif not isinstance(value, bool):\n                    cmd.append(str(value))\n            output = subprocess.run(cmd, capture_output=True)\n\n            # any return codes besides 0 mean there\n            # was an error thrown, so print the error\n            # then throw one of our own to get off this\n            # crazy thing\n            if output.returncode:\n                print(\n                    f\"Process failed with returncode {output.returncode}\"\n                )\n                print(output.stderr.decode(\"utf-8\"))\n                print(output.stdout.decode(\"utf-8\"))\n                raise RuntimeError\n            output = output.stdout.decode(\"utf-8\")\n\n            if self.callback:\n                return self.callback(output, *args, **kwargs)\n            return output\n        return wrapper","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def script_callback(output, *args, **kwargs):\n    print(output)\n\n    model_dir = kwargs[\"model_dir\"]\n    predictions_path = os.path.join(\n        model_dir, \"predictions.csv\"\n    )\n    predictions = pd.read_csv(predictions_path, index_col=\"image_id\")\n    tf.io.gfile.remove(predictions_path)\n\n    with open(os.path.join(model_dir, \"history.json\"), \"r\") as f:\n        history = json.load(f)\n    return history, predictions\n\n\n@scriptify(wrap=True, callback=script_callback)\ndef fit_a_model(\n    run_params: typing.Union[str, RunParams],\n    train_files: typing.List[str],\n    valid_files: typing.List[str],\n    model: typing.Optional[str] = None,\n    loss_function: typing.Optional[typing.Callable] = None,\n    metrics: typing.Optional[typing.List] = None,\n    model_dir: typing.Optional[str] = None,\n    initial_epoch: int = 0\n) -> typing.Tuple[tf.keras.callbacks.History, pd.DataFrame]:\n    \"\"\"\n    Do a complete training run on the provided model\n    \"\"\"\n    if isinstance(run_params, str):\n        run_params = RunParams.from_json(run_params)\n\n    # start by defining out datasets\n    train_dataset = make_dataset(\n        train_files,\n        run_params.batch_size,\n        downscale=run_params.downscale\n    )\n\n    # include filenames with valid dataset so\n    # that we can use it later for prediction.\n    # During training, we'll map a function\n    # onto it to strip out the filename\n    valid_dataset = make_dataset(\n        valid_files,\n        run_params.batch_size*4,\n        downscale=run_params.downscale,\n        shuffle=False,\n        return_filenames=True\n    )\n    valid_map = lambda X, y: (X[\"raw_input\"], y)\n\n    if initial_epoch > 0 or isinstance(model, str):\n        # we've either done some training already,\n        # or we explicitly pointed to a saved model\n        # to load. In either case, try to find a\n        # model to load and load it\n        if model is None and model_dir is not None:\n            model = os.path.join(model_dir, \"model\")\n        elif model_dir is None:\n            # we didn't specify a model or a place to look\n            # for one, so error out\n            raise ValueError(\"Must specify model directory!\")\n        model = tf.keras.models.load_model(model)\n\n    elif model is None:\n        # we didn't indicate that there's a model to\n        # load, so instantiate one to begin with\n        model = generic_cnn(\n            run_params.backbone,\n            run_params.input_shape,\n            classifier_hidden_dims=run_params.hidden_dims,\n            crop_factor=run_params.crop_factor,\n            contrast=run_params.contrast,\n            rotation=run_params.rotation,\n            zoom=run_params.zoom,\n            blue_noise=run_params.blue_noise\n        )\n\n        lr_schedule = tf.keras.experimental.CosineDecay(\n            run_params.learning_rate,\n            decay_steps=run_params.decay_steps\n        )\n\n        # compile the model's training step\n        if loss_function is None:\n            loss_function = tf.keras.losses.CategoricalCrossentropy(\n                label_smoothing=run_params.label_smoothing\n            )\n\n        metrics = metrics or [\"accuracy\"]\n        optimizer = tf.keras.optimizers.Adam(learning_rate=lr_schedule)\n        model.compile(optimizer, loss_function, metrics=metrics)\n\n    # add callback for early stopping\n    # configure to load back in the best\n    # performing weights at the end\n    callbacks = [HardBaselineEarlyStopping(\n        monitor=\"val_accuracy\",\n        patience=run_params.patience,\n        baseline=0.6, # if we can't do better than this, stop\n        restore_best_weights=True\n    )]\n\n    # fit the model!\n    history = model.fit(\n        train_dataset,\n        validation_data=valid_dataset.map(valid_map),\n        epochs=run_params.epochs,\n        callbacks=callbacks,\n        initial_epoch=initial_epoch,\n        verbose=2\n    )\n\n    # predict on the validation set\n    # TODO: use `training=False`? Or better to\n    # implement iteration + averaging in\n    # `predict_on_dataset`. If data loading\n    # is bottleneck (which is likely), this\n    # is probably most efficient\n    predictions = predict_on_dataset(\n        model, valid_dataset, training=False\n    )\n\n    # save weights if given the option\n    if model_dir is not None:\n        tf.io.gfile.makedirs(model_dir)\n        model.save(os.path.join(model_dir, \"model\"))\n\n    # return the training history and validation predictions\n    return history, predictions","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now create a command line argument parser for running this function as a script.","metadata":{}},{"cell_type":"code","source":"# last cell\nif __name__ == \"__replace_me__\":\n    import argparse\n    parser = argparse.ArgumentParser()\n    parser.add_argument(\n        \"--run-params\",\n        type=str,\n        required=True\n    )\n    parser.add_argument(\n        \"--train-files\",\n        type=str,\n        nargs=\"+\",\n        required=True\n    )\n    parser.add_argument(\n        \"--valid-files\",\n        type=str,\n        nargs=\"+\",\n        required=True\n    )\n    parser.add_argument(\n        \"--model-dir\",\n        type=str,\n        required=True\n    )\n    parser.add_argument(\n        \"--initial-epoch\",\n        type=int,\n        default=0\n    )\n    flags = parser.parse_args()\n    history, predictions = fit_a_model(**vars(flags))\n\n    predictions.to_csv(\n        os.path.join(flags.model_dir, \"predictions.csv\"),\n        index_label=\"image_id\"\n    )\n\n    # save out the history dict so that we can\n    # use the validation score information to\n    # do ensembling later\n    history_path = os.path.join(flags.model_dir, \"history.json\")\n    with open(history_path, \"w\") as f:\n        json.dump(history.history, f)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, write all the lines so far into a script called `main.py`","metadata":{}},{"cell_type":"code","source":"# loop through all the code cells above,\n# and dump their text content into a string\nscript, i = \"\", 1\nwhile i > 0:\n    try:\n        contents = _ih[i]\n    except IndexError:\n        break\n\n    # make sure we don't run `fit_a_model` in\n    # another subprocess\n    contents = contents.replace(\"wrap=True\", \"wrap=False\")\n\n    if contents.startswith(\"# last cell\"):\n        contents = contents.replace(\"__replace_me__\", \"__main__\")\n\n        # get rid of overly explanatory comments,\n        # add in blank last line for style,\n        # and reset `i` to break\n        contents = re.sub(\"(?<! )# .+\\n\", \"\", contents)\n        contents += \"\\n\"\n        i = -1\n    else:\n        # add in lines to break up functions\n        contents += \"\\n\\n\\n\"\n\n    script += contents\n    i += 1\n\n# write that string to a file\nwith open(\"main.py\", \"w\") as f:\n    f.write(script)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Cross Validation\nNow that I have my script ready, I can call `fit_a_model` from my cross validation function in exactly the same way I might if I was calling the function itself, only now it will be run as a subprocess in the bacgkround! Again, I'm not going to tell you that this is the best solution in the world, or that it shouldn't function as the world's best advertisement for PyTorch usage on Kaggle, but we got the functionality we needed and did it with tools that we can use in any competition going forward, which frankly I'll take if you'll look through the version history of this notebook and see how many OOM results I got.","metadata":{}},{"cell_type":"code","source":"def do_a_cross_validation(\n    run_params: RunParams,\n    files: typing.Optional[typing.List[str]] = None,\n    n_splits: int = 5,\n    run_name: typing.Optional[str] = None,\n    initial_epoch: typing.Optional[int] = None\n) -> typing.Tuple[str, pd.DataFrame, float]:\n    \"\"\"\n    Cross validate a given architecture and hyperparameter\n    settings over the entire dataset.\n    \"\"\"\n    if run_name is None:\n        run_name = str(run_params) + f\"-n_splits={n_splits}\"\n    run_dir = output_path(run_name)\n    tf.io.gfile.makedirs(run_dir)\n\n    config_path = os.path.join(run_dir, \"config.json\")\n    run_params.to_json(config_path)\n    # TODO: initialize a model here and save its weights\n    # to run dir if `initial_epoch == 0`, that way\n    # everything is initialized the same?\n\n    # cross validate the best way we can with tfrecords:\n    # by cross validating on the files themselves. Obviously\n    # this limits the randomness, but it makes iteration faster,\n    # and that's a tradeoff I'm prepared to make\n    files = files or glob.glob(data_path(\"train_tfrecords\", \"*.tfrec\"))\n    kfold = KFold(n_splits).split(files)\n\n    # loop through our kfold and collect predictions on\n    # the validation set\n    all_predictions = []\n    for fold, (train_idx, valid_idx) in enumerate(kfold):\n        train_files = [files[i] for i in train_idx]\n        valid_files = [files[i] for i in valid_idx]\n\n        msg = \"Fold {}/{}\".format(fold+1, n_splits)\n        box_print(msg, 7)\n\n        fold_dir = os.path.join(run_dir, str(fold))\n        tf.io.gfile.makedirs(fold_dir)\n\n        history, predictions = fit_a_model(\n            run_params=config_path,\n            train_files=train_files,\n            valid_files=valid_files,\n            model_dir=fold_dir,\n            initial_epoch=initial_epoch\n        )\n\n        # collect the validation data predictions\n        all_predictions.append(predictions)\n\n    # concatenate all of the prediction dataframes\n    # and compute a total accuracy score from the argmax\n    predictions = pd.concat(all_predictions, axis=0)\n    predicted_labels = predictions.drop(\"label\", axis=1).values.argmax(1)\n    accuracy = (predictions.label == predicted_labels).mean()\n    print(\"\\nMean accuracy: {}\\n\".format(accuracy))\n\n    # export the predictions for analysis after the fact\n    predictions.to_csv(\n        os.path.join(run_dir, \"predictions.csv\"),\n        index_label=\"image_id\"\n    )\n\n    run_params.to_json(os.path.join(run_dir, \"params.json\"))\n    return run_name, predictions, accuracy","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Make a Submission\nOnce we have a model (or models) trained up, use each of their cross validated weights to generate a submission file, optionally normalizing by each fold's validation accuracy.","metadata":{}},{"cell_type":"code","source":"def make_a_submission(\n    model_names: typing.Union[str, typing.List[str]],\n    output_file: typing.Optional[str] = None,\n    weight: bool = True,\n    batch_size: int = 128,\n    iterations: int = 5\n) -> pd.DataFrame:\n    if isinstance(model_names, str):\n        model_names = [model_names]\n\n    predictions = []\n    test_files = glob.glob(data_path(\"test_tfrecords\", \"*.tfrec\"))\n    for model_name in model_names:\n        model_dir = output_path(model_name)\n\n        # try to load some run params from either\n        # a saved json or the model name string if\n        # it doesn't exist. If both fail, error out\n        try:\n            run_params = RunParams.from_json(\n                os.path.join(model_dir, \"params.json\")\n            )\n        except FileNotFoundError:\n            try:\n                run_params = RunParams.from_string(model_name)\n            except Exception as e:\n                raise ValueError(\n                    \"Can't find config to load from\"\n                ) from e\n\n        versions = next(tf.io.gfile.walk(model_dir))[1]\n        if weight:\n            # build weightings using best validation accuracy\n            weights = []\n            for version in versions:\n                history_path = os.path.join(\n                    model_dir, version, \"history.json\"\n                )\n                # try to load accuracies from saved histories\n                try:\n                    with open(history_path, \"r\") as f:\n                        history = json.load(f)\n                    weights.append(max(history[\"val_accuracy\"]))\n                except FileNotFoundError:\n                    weights.append(None)\n\n            if all([weight is None for weight in weights]):\n                # if none of the versions have histories,\n                # then ignore and weight everything equally\n                weights = [1. for _ in versions]\n            elif any([weight is None for weight in weights]):\n                # if some of them do, fill with mean value\n                valid_weights = [\n                    weight for weight in weights if weight is not None\n                ]\n                mean_value = sum(valid_weights) / len(valid_weights)\n                weights = [weight or mean_value for weight in weights]\n\n            versions = dict(zip(versions, weights))\n        else:\n            # otherwise weight everything the same\n            versions = {version: 1. for version in versions}\n\n        # instantiate dataset here since we need the downscale\n        test_dataset = make_dataset(\n            test_files,\n            run_params.batch_size*4,\n            downscale=run_params.downscale,\n            shuffle=False,\n            return_filenames=True,\n            label=False\n        )\n\n        model_predictions = 0\n        for version, weight in versions.items():\n            # iterate through each version of the model\n            # load its weights in\n            model = tf.keras.models.load_model(\n                os.path.join(model_dir, version, \"model\")\n            )\n\n            version_predictions = 0\n            for i in range(iterations):\n                # for each version, do multiple predictions and\n                # average them (I need to figure out if the\n                # preprocessing if random at inference time)\n                version_predictions += predict_on_dataset(\n                    model, test_dataset\n                )\n            version_predictions /= iterations\n\n            # weight these average predictions and add them\n            # to the model predictions\n            model_predictions += version_predictions*weight\n\n        # divide model predictions by the sum of the weights\n        # and add to the predictions for _all_ models\n        model_predictions /= sum(versions.values())\n        predictions.append(model_predictions)\n\n    # finally do a simple averaging over each model's predictions\n    final_predictions = 0\n    for prediction in predictions:\n        final_predictions += prediction\n    final_predictions /= len(predictions)\n\n    # take the label as the argmax\n    final_predictions[\"label\"] = final_predictions.values.argmax(axis=1)\n    final_predictions = final_predictions[\"label\"].reset_index()\n\n    output_file = output_file or \"submission.csv\"\n    final_predictions.to_csv(output_file, index=False)\n    return final_predictions","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Hyperparameter Searching\nAnd as a final layer, we'll abstract one more level to build a function to search through a few hyperparameters for us, that way we can press play, walk away, and peruse our outputs when they're ready. We'll use the `kerastuner` library to do implement the hyperparameter sampling, but build our own (approximate) implementation of a single [Hyperband](https://arxiv.org/abs/1603.06560) bracket, since the out-of-the-box `kerastuner` implementation won't play well with our cross-validation based approach (though it may down the line, and it has other nice HP optimization algorithms pre-built, so using its `HyperParameters` implementation will hopefully future-proof us). This is about all we can do with our current time constraints, so we son't do the normal search over `num_initial_points`/compute time.","metadata":{}},{"cell_type":"code","source":"def _sort_result_ids(results):\n    key = lambda item: item[1]\n    return [id for id, acc in sorted(results.items(), key=key)]\n\n\n@attr.s(auto_attribs=True)\nclass HyperbandOptimizer:\n    \"\"\"\n    Object for running Hyperband hyperparameter\n    optimization searches\n\n    :param run_params: Global `RunParams` to use for each\n        HP search trial. Note that each sampled hyperparameter\n        value will be subbed in for the corresponding value in\n        the `RunParams`, and the `epochs` value will be subbed\n        out at for each value in `epochs` during each round\n    :param files: Files to use for the HP search, either provided\n        as an explicit list of filename strings, or as an `int`\n        that will sample that many filenames for all the available\n        files. If left as `None`, all available files will be used.\n    :param n_splits: number of cross validation splits to use. If\n        greater than the number of files in `files`, the number\n        of files specified will be used.\n    :param num_initial_points: number of initial trials to use\n    :param r: base number of epochs for which to run modelsS\n    :param eta: fraction by which to pair down trials after\n        each round, see the paper for more\n    \"\"\"\n    run_params: RunParams\n    files: typing.Optional[typing.Union[typing.List[str], int]] = None\n    n_splits: int = 4\n    num_initial_points: int = 18\n    r: int = 1\n    eta: int = 3\n\n    # could also be achieved via converters\n    def __attrs_post_init__(self):\n        if isinstance(self.files, int):\n            # sample the given number of files if an\n            # integer was provided\n            self.files = random.sample(\n                glob.glob(data_path(\"train_tfrecords\", \"*.tfrec\")),\n                k=self.files\n            )\n        if self.files is not None:\n            # if we have files explicitly specified, pare down\n            # the number of splits to match it, otherwise we\n            # can't properly cross-validate\n            self.n_splits = min(self.n_splits, len(self.files))\n        \n\n    def _run_a_round(\n        self,\n        trials: typing.Dict[str, dict],\n        initial_epoch: int\n    ) -> typing.Dict[str, float]:\n        \"\"\"\n        Run a single round of hyperband optimization\n        using the given trials, proceeding from the\n        indicated epoch.\n\n        :param trials: `dict` mapping from trial ID strings\n            to `dict`s mapping from hyperparameter names to\n            sampled values for a given trial\n        :param initial_epoch: The epoch from which to begin\n            training for each fold in each trial\n        \"\"\"\n        msg = \"Running search on {} models for {} epochs\".format(\n            len(trials),\n            self.run_params.epochs - initial_epoch\n        )\n        box_print(msg, 2)\n\n        num_to_keep = max(1, len(trials) // self.eta)\n        results = {}\n        for trial_id, values in trials.items():\n            for hp_name, value in values.items():\n                setattr(self.run_params, hp_name, value)\n\n            results[trial_id] = do_a_cross_validation(\n                self.run_params,\n                files=self.files,\n                n_splits=self.n_splits,\n                run_name=trial_id,\n                initial_epoch=initial_epoch\n            )[-1]\n\n            # due to disk constraints, delete our worst\n            # performing trials up front\n            if len(results) > num_to_keep:\n                # iterate through all our results and find the\n                # first bad one that still has a directory\n                for id in _sort_result_ids(results):\n                    if tf.io.gfile.exists(output_path(id)):\n                        tf.io.gfile.rmtree(output_path(id))\n                        break\n\n        trial_ids = _sort_result_ids(results)[::-1]\n        good_ids = trial_ids[:num_to_keep]\n        bad_ids = trial_ids[num_to_keep:]\n\n        print(\n            \"Removing ids:\\n\\t{}\\n\".format(\"\\n\\t\".join(bad_ids))\n        )\n\n        N = min(4, len(good_ids))\n        msg = \"Top {} trials for round of {}:\".format(N, len(trials))\n        box_print(msg, 2)\n\n        for id in good_ids[:N]:\n            pct = results[id]*100\n            print(f\"Validation Accuracy for Trial {id}: {pct:0.2f}%\")\n            for hp, value in trials[id].items():\n                print(f\"\\t{hp}: {value}\")\n        print(\"\\n\\n\")\n\n        return {id: results[id] for id in good_ids}\n\n    def fit(\n        self, hyperparameters: kt.HyperParameters\n    ) -> typing.Dict[str, float]:\n        \"\"\"\n        Run a Hyperband optimization over the given\n        hyperparameter space using the config parameters\n\n        :param hyperparameters: The space of hyperparameters\n            from which to sample and over which to optimize\n        \"\"\"\n        # initialize a bunch of search points, give\n        # them a random indentifier\n        trials = {}\n        hp_space = hyperparameters.space\n        for i in range(self.num_initial_points):\n            params = {p.name: p.random_sample() for p in hp_space}\n            trial_id = hex(hash(str(params))).split(\"x\")[1]\n            trials[trial_id] = params\n\n        # run the search\n        initial_epoch, n = 0, 0\n        max_epochs = self.run_params.epochs\n        while len(trials) > 1:\n            epochs = self.r*self.eta**n\n            if (len(trials) // self.eta) <= 1:\n                # if this is going to be the last round,\n                # let it run for the max time\n                epochs = max(max_epochs, epochs)\n\n            self.run_params.epochs = epochs\n            results = self._run_a_round(trials, initial_epoch)\n\n            # only keep the returned trial ids\n            trials = {id: trials[id] for id in results}\n            initial_epoch = epochs\n            n += 1\n\n        # return the results from the last round\n        # to know which model ID to use for our\n        # prediction (and estimate our score)\n        return results\n\n\ndef search_and_submit(\n    hyperparameters: kt.HyperParameters,\n    hp_optimizer: HyperbandOptimizer,\n    output_file: str = \"submission.csv\",\n    iterations: int = 10\n) -> typing.Tuple[pd.DataFrame, str, float]:\n    results = hp_optimizer.fit(hyperparameters)\n    trial_id = list(results.keys())[0]\n    test_predictions = make_a_submission(\n        trial_id,\n        weight=True,\n        output_file=output_file,\n        iterations=iterations\n    )\n    return test_predictions, trial_id, results[trial_id]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Putting It All Together\nNow that we have all the tools we need, at multiple levels of abstraction, all that remains is to run our hyperparameter search, build the best model we can, and then create a submission from it.\n\nTo do this, we just specify a set of `RunParams` describing how we want individual model training runs to go, then specify the `HyperParameters` that we'll search over and change\nfor each run. Then it's as simple as using our `search_and_submit` function above to take\ncare of everything else. This way, running new tests with different backbones or different\noptimization parameters amounts to just changing lines in our config. Of course, it didn't come _free_ (as my ever-patient girlfriend who put up with my obsessive writing of this will attest), but it should now make further experimentation simple and robust (and hopefully we learned some things along the way).","metadata":{}},{"cell_type":"code","source":"params = RunParams(\n    backbone=\"res_net50_v2\",\n    hidden_dims=[1024],\n    batch_size=32,\n    learning_rate=1e-4,\n    downscale=2,\n    crop_factor=0.875,\n    decay_steps=100,\n    epochs=20,\n    patience=4,\n    label_smoothing=0.2\n)\n\nhp = kt.HyperParameters()\nhp.Float(\"learning_rate\", 5e-6, 5e-4, sampling=\"log\")\nhp.Int(\"decay_steps\", 50, 50000, sampling=\"log\")\nhp.Float(\"blue_noise\", 0.01, 0.3, sampling=\"log\")\n\nhp_optimizer = HyperbandOptimizer(\n    run_params=params\n)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if TEST:\n    # do a quick test up front to make\n    # sure the pipeline can run end-to-end before\n    # investing a bunch of time in it\n    test_params = params.copy()\n    test_params.epochs = 2\n    test_hp_optimizer = HyperbandOptimizer(\n        run_params=test_params,\n        files=4,\n        n_splits=4,\n        num_initial_points=4,\n        r=2,\n        eta=2\n    )\n\n    # make a fake submission then remove the\n    # model directory since we don't need it\n    test_test_predictions, test_trial_id, test_val_accuracy = \\\n        search_and_submit(hp, test_hp_optimizer, iterations=3)\n    tf.io.gfile.rmtree(output_path(test_trial_id))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if SUBMIT:\n    # the real function call we care about\n    test_predictions, trial_id, val_accuracy = search_and_submit(\n        hp, hp_optimizer, iterations=10\n    )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## TODOs\nAlways more to be done. The highest priority items, off the top of my head, include\n* Trying `efficient_net_b4` with a smaller batch size (64 crashes GPU memory)\n* Using `albumentations` for more random augmentation\n* At inference time, producing a prediction an _unaugmented_ version of the image, to averaged with the _averaged_ augmented test predictions (essentially giving it a (normalized) averaging weight of 0.5)\n* Switching to cosine decay for learning rate\n* Configuring TF with CPU at the start to show what individual functions do, and introducing a `# script ignore` comment at the start of these cells to keep them from going into the cross validation script.\n* Profiling the training to identify and eliminate bottlenecks. Since P100 GPUs won't benefit from mixed precision compute, the main thing to look at would be data loading bottlenecks, which could possibly be managed with DALI (though non-trivial to get working in this environment).","metadata":{}}]}