{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Stratified K-Fold Cross-Validation and Resnet34 with fast.ai","metadata":{}},{"cell_type":"markdown","source":"- The first notebook of this serie was a simple baseline using resnet34: [Fast Resnet34 with Fastai](https://www.kaggle.com/code/fmussari/fast-resnet34-with-fastai)  \n- In this notebook we are going to explore that same model, ensembling the trainings of 5 folds to see how much that improves the accuracy.\n  \n","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://drive.google.com/uc?export=view&id=1EucGY8cJYJiuAZHBp95UdS22zVeWvyjl\" width=\"500\">","metadata":{}},{"cell_type":"markdown","source":"## Acknowledgements\n\n- [Fast Resnet34 with Fastai](https://www.kaggle.com/code/fmussari/fast-resnet34-with-fastai)\n\n**fastai course:**\n- [Practical Deep Learning for Coders (a UQ collaboration with fast.ai)](https://itee.uq.edu.au/event/2022/practical-deep-learning-coders-uq-fastai)  \n\n**Jeremy's Notebook Series:**\n- [First Steps: Road to the Top, Part 1](https://www.kaggle.com/code/jhoward/first-steps-road-to-the-top-part-1)\n- [Small models: Road to the Top, Part 2](https://www.kaggle.com/code/jhoward/small-models-road-to-the-top-part-2)\n- [Scaling Up: Road to the Top, Part 3](https://www.kaggle.com/code/jhoward/scaling-up-road-to-the-top-part-3)\n- [Multi-target: Road to the Top, Part 4](https://www.kaggle.com/code/jhoward/multi-target-road-to-the-top-part-4)","metadata":{}},{"cell_type":"markdown","source":"## K-Fold","metadata":{}},{"cell_type":"markdown","source":"- K-Fold is a technique in which data is divided into K parts. In this case we are going to use K=5.\n- We will do 5 trainings, each one with a different validation set (in each experiment we are going to take one fold for validation and the other four for training).\n- This way we'll end up with 5 models trained in slightly different data, but with completely different validations.\n- We don't know wich of them is the best model, but we could assume that taking the mean of them would be the best for generalization.\n- We are going to use [KFold from sklearn](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.KFold.html).","metadata":{}},{"cell_type":"markdown","source":"## Installing the libraries\n","metadata":{}},{"cell_type":"code","source":"# fastkaggle allows you to work locally and then submit the results and notebook to Kaggle\n\ntry: import fastkaggle\n\nexcept ModuleNotFoundError:\n    !pip install -Uq fastkaggle\n\nfrom fastkaggle import *\nfrom sklearn.model_selection import KFold\nimport plotly.express as px\nfrom datetime import datetime as dt","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"competition = 'paddy-disease-classification'\npath = setup_comp(competition, install='fastai \"timm>=0.6.2.dev0\"')\n\nfrom fastai.vision.all import *\n#from scipy.special import softmax, log_softmax\nimport gc","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Setting data paths","metadata":{}},{"cell_type":"code","source":"# train images\ntrain_path = path / 'train_images'\ntrain_files = get_image_files(train_path)\n\n# test images\ntest_path = path/'test_images'\ntest_files = get_image_files(test_path).sorted()\n\n# sample submission\nsample_submission = pd.read_csv(path/'sample_submission.csv')\n\n# train labels\ntrain_df = pd.read_csv(path / 'train.csv')\n","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### target distribution","metadata":{}},{"cell_type":"code","source":"train_df.label.value_counts() * 100 / len(train_df)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Stratified Folding\n- If we apply sklearn KFold to all the dataset, we could end up with 5 folds that don't have the same distributions by targets as our full dataset.\n- There is a sklearn function called [StratifiedKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html#sklearn.model_selection.StratifiedKFold) to assure that each fold preserves the percentage of samples for each target class.\n- Or, as we are going to do here, we can apply K-Fold individually to each of the labels. If we divide each class in 5 randomly, we are sure that at the end, when we join the pieces, we are going to have the same distribution.\n- To keep track of which image belongs to what fold, we can create a field called `kfold` and initializa it with","metadata":{}},{"cell_type":"code","source":"train_df['kfold'] = -1\n\n# Number of Splits\nn_folds=5\nreversed = False\n\n# For each label we are going to create n_folds folds\nfor label in train_df.label.unique():\n    \n    # Assign fold number from 0 to n_folds or from n_folds to 0\n    # Because KFold assigns less data for last fold (we assign it to fold 0 or n_folds)\n    folds = list(range(n_folds))\n    if reversed: folds.reverse()\n    \n    kf = KFold(n_splits=n_folds, random_state=42, shuffle=True)\n    \n    # Indices for each label\n    label_idxs = train_df[train_df.label==label].index\n    \n    # Creating folds for those indices\n    kf.get_n_splits(label_idxs)\n\n    for _, valid_index in kf.split(label_idxs):\n\n        actual_fold = folds.pop(0)\n        df_index = label_idxs[valid_index]\n        train_df.loc[df_index, 'kfold'] = actual_fold\n    reversed = not reversed\n        ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.sample(5)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### plot distributions by fold","metadata":{}},{"cell_type":"code","source":"df = train_df.groupby(['label', 'kfold']).size().reset_index()\ndf.columns = ['label', 'kfold', 'count']\n#df.kfold = df.kfold.astype('str')\n\nfig = px.bar(\n    df, x=\"kfold\", y=\"count\",\n    color='label', barmode='group',\n    height=400\n)\nfig.show()","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Custom Split Function\n- In the first notebook we passed the following splitter to the `DataBlock`:\n```\nsplitter=RandomSplitter(0.2, seed=42),\n```\n- In this case we are going to use a custom function as splitter.","metadata":{}},{"cell_type":"markdown","source":"### FuncSplitter\n- fast.ai `FuncSplitter` needs a function to be passed that returns `True` for items that belongs to validation set.\n- So we are going to create a list of dictionaries (one for each fold) containing that information.\n","metadata":{}},{"cell_type":"code","source":"img2valid = []\n\nfor fold in range(n_folds):\n    train_df['is_valid'] = False\n    idxs = train_df[train_df.kfold == fold].index\n    train_df.loc[idxs, 'is_valid'] = True\n    \n    img2valid.append({ r.image_id: r.is_valid for _, r in train_df.iterrows() })","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- There we are, a list of dictionaries that for each fold, has `True` for each validation image, and now it is easy to retrieve, for each fold, if an image belong to the validation set or not.","metadata":{}},{"cell_type":"markdown","source":"## Dataloaders for fastai training\n","metadata":{}},{"cell_type":"code","source":"def get_datablock(i_fold, size, item_tfms, accum):\n    \n    def get_split(p):\n        # For each fold, return if an image is in valid set or not\n        return img2valid[i_fold][p.name]\n    \n    dblock = DataBlock(\n        blocks=(ImageBlock, CategoryBlock),\n        get_items=get_image_files,\n        get_y=parent_label,\n        # Custom Splitter\n        splitter = FuncSplitter(get_split),\n        item_tfms=item_tfms,\n        batch_tfms=aug_transforms(size=size, min_scale=0.75)\n    )\n    return dblock.dataloaders(train_path, bs=64//accum)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training Function","metadata":{}},{"cell_type":"code","source":"def train(i_fold, arch, size=224, item_tfms=Resize(480, method='squish'), accum=1, epochs=16, lr=0.005):\n    \n    dls = get_datablock(i_fold=i_fold, size=size, item_tfms=item_tfms, accum=accum)\n    print('- First 5 validation images:')\n    print([each.name for each in dls.valid.items[:5]])\n    \n    cbs = GradientAccumulation(64) if accum!=1 else []\n    \n    # Force torchvision models instead of TIMM, when possible\n    try: arch = eval(arch)\n    except: arch = arch\n        \n    learn = vision_learner(dls, arch, metrics=error_rate, cbs=cbs).to_fp16()\n    print('- Fine Tuning')\n    learn.fine_tune(epochs, lr)\n    #print('- Getting predictions')\n    #probs, _ = learn.get_preds(dl=dls.test_dl(test_files))\n    probs = None # Return only tta_preds\n    print('- Getting tta_predictions')\n    preds, _ = learn.tta(dl=dls.test_dl(test_files))\n    \n    return probs, preds, dls.vocab","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Running the Model(s) on Selected Fold\n- In this notebook, as in ([Fast Resnet34 with Fastai](https://www.kaggle.com/code/fmussari/fast-resnet34-with-fastai)) notebook, we are going to experiment with resnet34.\n- You can create a copy of this notebook and, just as Jeremy did in his notebook ([Scaling Up: Road to the Top, Part 3](https://www.kaggle.com/code/jhoward/scaling-up-road-to-the-top-part-3)), try different models with control over the validation sets you want to use for each model or experiment. Remeber to set `accum` according to model size and available GPU memory.","metadata":{}},{"cell_type":"markdown","source":"### resnet34 in each fold","metadata":{}},{"cell_type":"code","source":"models = {\n    'resnet34': {\n        (0, Resize(480, method='squish'), 224),\n        (1, Resize(480, method='squish'), 224),\n        (2, Resize(480, method='squish'), 224),\n        (3, Resize(480, method='squish'), 224),\n        (4, Resize(480, method='squish'), 224)\n    }\n    \n}","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = []\ntta_predictions = []","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### run the experiments","metadata":{}},{"cell_type":"code","source":"exp = 1\nfor arch, details in models.items():\n    \n    for i_fold, item, size in details:\n        print('////'*10)\n        print('---Experiment', exp, '--', arch)\n        print('fold: ', i_fold)\n        print(item.name)\n        \n        preds, tta_preds, vocab = train(i_fold, arch, size, item_tfms=item, accum=1, epochs=20, lr=0.005)\n        \n        predictions.append(preds)\n        tta_predictions.append(tta_preds)\n        \n        now = dt.now().strftime(\"%Y%m%d\")\n        filename = f'{now}-exp{exp}-{arch}-Fold{i_fold}.csv'\n        print(f'Saving {filename}')\n        \n        sample_submission.label = vocab[tta_preds.argmax(axis=1)]\n        sample_submission.to_csv(filename, index=False)\n        \n        gc.collect()\n        torch.cuda.empty_cache()\n        exp += 1","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Ensembling","metadata":{}},{"cell_type":"code","source":"[each.shape for each in tta_predictions]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"avg_tta_predictions = torch.stack(tta_predictions).mean(0)\navg_tta_predictions.shape","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Final Submission","metadata":{}},{"cell_type":"code","source":"vocab[avg_tta_predictions.argmax(dim=1)]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission.label = vocab[avg_tta_predictions.argmax(dim=1)]\n\nsample_submission.to_csv('submission.csv', index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusions","metadata":{}},{"cell_type":"markdown","source":"- Maybe it is more useful to try different models with different random splits as Jeremy did in his [Scaling Up: Road to the Top, Part 3](https://www.kaggle.com/code/jhoward/scaling-up-road-to-the-top-part-3), instead of using cross-validation (different validations for each training).\n- The results when running it locally were the following for 16 epochs: ","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://drive.google.com/uc?export=view&id=1EucGY8cJYJiuAZHBp95UdS22zVeWvyjl\" width=\"500\">","metadata":{}},{"cell_type":"markdown","source":"- The ensembled model had a score of **0.98615** which is better than the best score for an individual fold (0.98269).","metadata":{}},{"cell_type":"code","source":"# Pushing the notebook from my home PC to Kaggle\n\nif not iskaggle:\n    push_notebook(\n        'fmussari', \n        'Stratified Cross-Validation, Resnet34 & Fastai',\n        title='Stratified Cross-Validation, Resnet34 & Fastai',\n        file='2022-07. Cross-validation and Resnet34 with Fastai [Submission].ipynb',\n        competition=competition, \n        private=False, \n        gpu=True\n    )","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}