{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":6799,"databundleVersionId":4225553,"sourceType":"competition"},{"sourceId":7556785,"sourceType":"datasetVersion","datasetId":4400759},{"sourceId":7582952,"sourceType":"datasetVersion","datasetId":4413931}],"dockerImageVersionId":30648,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Searching for Hardware Aware Neural Networks on ImageNet for FPGA in MobilenetV3 Search Space\n\nWe need to estimate the accuracy of samples in our search space because it requires significant amount of time to train each subnet for evaluating its perfomace. This slows down the searching process. Accuracy peredictory MLP that was trained on the (arch, accu) dataset in our search space is used to speed up the search process as we won't have to train the network at all. The search is perfomed on the Imagenet Dataset. \n\nWe have also build a latency table by manualy deploying all the unique blocks of our search space on ULTRA96v2 FPGA board this will be utilized to estimate latency of our networks.\n\n\n#### IDEA\nIntroducing a new parameter by combining the model size(params), mac ops , and latency called efficient_arthemetic_intensity which is defined   \n\nefficient_arthemetic_intensity = macs/(model_size*latency) ops/(byte*ms)  \n\nTrying to do three things maximizing the mac operations perfomed by model which represents the learning capacity\nof model while minimizing the number of paramters which is the size of model and also minimizing the latency \non target hardware thus finding a model that has performs maximum calculation while being smaller and faster.\n\n\n### Overview\nEvery architecutre configuration is represented in dictionary format.  \nAccuracy predictor predicts accuracy based on this configuration.  \nLatency Estimator also uses this configuration to estimate the latency.   \n","metadata":{"_cell_guid":"b72d2cc2-6903-4a1a-b966-dd65184120a5","_uuid":"214f39c3-e9b3-4d4a-9191-7fbccd0da632","trusted":true}},{"cell_type":"markdown","source":"## 1. Preparation\nLet's first install all the required packages:","metadata":{"_cell_guid":"7ff1f4fe-c4e8-464e-bcd0-f05d4a2fb091","_uuid":"ca674990-a38b-40ff-9e89-cc0242feac5a","trusted":true}},{"cell_type":"code","source":"# For kaggle\n!rm -r /kaggle/working/Evolutionary-Neural-Architectural-Search-for-FPGAs /kaggle/working/ofa /kaggle/working/viz /kaggle/working/search_space_blocks /kaggle/working/blocks /kaggle/working/figures\n\n!pip install thop \n! pip install gdown\n!pip install shutil\n!pip install graphviz\n! pip install torch-summary \n! git clone --branch for_kaggle https://github.com/amitpant7/Evolutionary-Neural-Architectural-Search-for-FPGAs.git\n! mv -f /kaggle/working/Evolutionary-Neural-Architectural-Search-for-FPGAs/* /kaggle/working\n! rm -r Evolutionary-Neural-Architectural-Search-for-FPGAs","metadata":{"_cell_guid":"5f7d2c59-4621-45ce-bf03-40e532baaf47","_uuid":"2541a397-5182-423f-a65e-a83893b1a12a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-06-01T15:22:41.910979Z","iopub.execute_input":"2024-06-01T15:22:41.911452Z","iopub.status.idle":"2024-06-01T15:24:23.016548Z","shell.execute_reply.started":"2024-06-01T15:22:41.911399Z","shell.execute_reply":"2024-06-01T15:24:23.015005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install thop \n! pip install gdown\n!pip install shutil\n!pip install graphviz\n! pip install torch-summary","metadata":{"_cell_guid":"e8f24185-24c8-4f69-98fc-2234a3274353","_uuid":"e0f9376e-390f-47e2-8026-06088f0c7457","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-06-01T15:24:23.019183Z","iopub.execute_input":"2024-06-01T15:24:23.019634Z","iopub.status.idle":"2024-06-01T15:25:28.599024Z","shell.execute_reply.started":"2024-06-01T15:24:23.019585Z","shell.execute_reply":"2024-06-01T15:25:28.597659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Installing PyTorch...')\n! pip install torch \nprint('Installing torchvision...')\n! pip install torchvision\nprint('Installing numpy...')\n! pip install numpy \n# thop is a package for FLOPs computing.\nprint('Installing thop (FLOPs counter) ...')\n! pip install thop \n# ofa is a package containing training code, pretrained specialized models and inference code for the once-for-all networks.\n# print('Installing OFA...')\n# ! pip install ofa \n# tqdm is a package for displaying a progress bar.\nprint('Installing tqdm (progress bar) ...')\n! pip install tqdm \nprint('Installing matplotlib...')\n! pip install matplotlib \n! pip install torch-summary \n\nprint('All required packages have been successfully installed!')","metadata":{"_cell_guid":"51e2d4b2-f708-426b-8006-f93ac6f1de51","_uuid":"c79b023d-9eb3-4fd9-87d4-624f89997d7d","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-06-01T15:25:28.601033Z","iopub.execute_input":"2024-06-01T15:25:28.601462Z","iopub.status.idle":"2024-06-01T15:27:18.140661Z","shell.execute_reply.started":"2024-06-01T15:25:28.601418Z","shell.execute_reply":"2024-06-01T15:27:18.13894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, we can import the packages used in this tutorial:","metadata":{"_cell_guid":"f552a9b7-c47b-4998-9e7a-376e85cc8737","_uuid":"bd7c56c3-f0d5-4cc1-886f-80adb0b34918","trusted":true}},{"cell_type":"code","source":"import os\nimport torch\nimport torch.nn as nn\nfrom torchvision import transforms, datasets\nimport numpy as np\nimport time\nimport random\nimport shutil\nimport math\nfrom PIL import Image\nimport copy\nfrom matplotlib import pyplot as plt\nfrom torchsummary import summary\n\nfrom ofa.model_zoo import ofa_net\nfrom ofa.utils import download_url\n\nfrom ofa.accuracy_predictor import AccuracyPredictor\nfrom ofa.flops_table import ArthIntTable\n\nfrom ofa.evolution_finder import EvolutionFinder\nfrom ofa.imagenet_eval_helper import evaluate_ofa_subnet, evaluate_ofa_specialized\nfrom ofa.imagenet_classification.elastic_nn.networks.ofa_mbv3 import OFAMobileNetV3\n\nfrom ofa.utils.arch_visualization_helper import draw_arch\n\nfrom tqdm import tqdm\n\n# set random seed\nrandom_seed = 1\nrandom.seed(random_seed)\nnp.random.seed(random_seed)\ntorch.manual_seed(random_seed)\nprint('Successfully imported all packages and configured random seed to %d!'%random_seed)","metadata":{"_cell_guid":"ed77b74d-713b-40d2-9248-96d4a9556ad4","_uuid":"ef7c9af0-0bcc-4cd2-827d-a2e219c22dc4","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-06-01T15:27:18.142994Z","iopub.execute_input":"2024-06-01T15:27:18.143488Z","iopub.status.idle":"2024-06-01T15:27:18.557779Z","shell.execute_reply.started":"2024-06-01T15:27:18.143441Z","shell.execute_reply":"2024-06-01T15:27:18.556405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now it's time to determine which device to use for neural network inference in the rest of this tutorial. If your machine is equipped with GPU(s), we will use the GPU by default. Otherwise, we will use the CPU.","metadata":{"_cell_guid":"0848ed15-078c-4fbb-90e8-0845550faec4","_uuid":"6cfbc39e-7dfd-4258-85a0-2b33f71a1aae","trusted":true}},{"cell_type":"code","source":"#os.environ['CUDA_VISIBLE_DEVICES'] = '0'\ncuda_available = torch.cuda.is_available()\nif cuda_available:\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.benchmark = True\n    torch.cuda.manual_seed(random_seed)\n    print('Using GPU.')\nelse:\n    print('Using CPU.')","metadata":{"_cell_guid":"d5958bd8-92cb-47f5-881b-aa17cf37efd3","_uuid":"dbf7d111-1be6-4b50-9e06-352221ee2dbc","collapsed":false,"jupyter":{"outputs_hidden":false},"pycharm":{"name":"#%%\n"},"execution":{"iopub.status.busy":"2024-06-01T15:27:18.561779Z","iopub.execute_input":"2024-06-01T15:27:18.562485Z","iopub.status.idle":"2024-06-01T15:27:18.570439Z","shell.execute_reply.started":"2024-06-01T15:27:18.562446Z","shell.execute_reply":"2024-06-01T15:27:18.569109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Dataset Preperation","metadata":{"_cell_guid":"8bbdc985-a6ba-4983-9df7-2105d9ba771b","_uuid":"4db057a3-8141-4ee3-98d7-675463897305","trusted":true}},{"cell_type":"markdown","source":"Now, let's build the ImageNet dataset and the corresponding dataloader. Notice that **if you are not in kaggle it will be skipped** since it will be very slow.\n\n\nWe will only use subset of ImageNet validation set which will contains 10,000 images for testing.","metadata":{"_cell_guid":"ebaa20c6-25ae-4f12-9982-1e7fdadee3bc","_uuid":"07c22216-4e65-4f25-9bd2-b1ac2f6d0e57","pycharm":{"name":"#%% md\n"},"trusted":true}},{"cell_type":"markdown","source":"**I will utilize subsets of imagenet for both validation and retraining.**","metadata":{"_cell_guid":"d5ff6436-4b58-45e7-8c39-1a4c00b171d3","_uuid":"919942ef-4c54-48ef-9f1f-aedcdaa044df","trusted":true}},{"cell_type":"code","source":"%%script echo skipping\nbatch_size=32\n\n#I will use a susbset of imagenetval of 10k images \nif cuda_available:\n    # path to the ImageNet dataset\n    # link --> https://www.kaggle.com/datasets/titericz/imagenet1k-val\n    \n    imagenet_data_path = '/kaggle/input/imagenet1k-subset-100k-train-and-10k-val/imagenet_subval'\n\n    # if 'imagenet_data_path' is empty, download a subset of ImageNet containing 2000 images (~250M) for test\n    if not os.path.isdir(imagenet_data_path):\n        print('%s is empty. Download a subset of ImageNet for test.' % imagenet_data_path)\n\n    print('The ImageNet dataset files are ready.')\nelse:\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')\n    \n    \n  \nif cuda_available:\n    # The following function build the data transforms for test\n    def build_val_transform(size):\n        return transforms.Compose([\n            transforms.Resize(int(math.ceil(size / 0.875))),\n            transforms.CenterCrop(size),\n            transforms.ToTensor(),\n            transforms.Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225]\n            ),\n        ])\n    \n    val_data = datasets.ImageFolder(\n            root=os.path.join(imagenet_data_path),\n            transform=build_val_transform(224)\n        )\n    \n\n    val_loader = torch.utils.data.DataLoader(\n        val_data,\n        batch_size=batch_size,  \n        shuffle = True,\n        num_workers=4,  \n        pin_memory=True,\n        drop_last=False,\n    )\n    print('The ImageNet dataloader is ready. Size : {}'.format(len(val_loader)*batch_size))\nelse:\n    data_loader = None\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')","metadata":{"_cell_guid":"dece162d-f096-48e6-a573-e69c2bebdba6","_uuid":"c8f20a68-77af-4a1e-aec2-6b4f968950d3","collapsed":false,"jupyter":{"outputs_hidden":false},"pycharm":{"name":"#%%\n"},"execution":{"iopub.status.busy":"2024-06-01T15:27:18.572Z","iopub.execute_input":"2024-06-01T15:27:18.572371Z","iopub.status.idle":"2024-06-01T15:27:18.603419Z","shell.execute_reply.started":"2024-06-01T15:27:18.572339Z","shell.execute_reply":"2024-06-01T15:27:18.601777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now you have configured the dataset. Let's build the dataloader for evaluation.\nAgain, this will be skipped if you are in a CPU environment.","metadata":{"_cell_guid":"59b9927f-0fc0-4e91-8ac4-ef6db0064ddc","_uuid":"1616ad0c-4e91-419d-9339-501a56f480fb","trusted":true}},{"cell_type":"markdown","source":"Lets evaluate our randomly sampled network on imagenet validation set","metadata":{"_cell_guid":"044134e6-a76c-49e8-bf41-9622c3e711e2","_uuid":"39890b5e-7794-4839-af6c-822eca266fe1","trusted":true}},{"cell_type":"code","source":"%%script echo skipping\nfrom ofa.utils.common_tools import *\ndef evaluate_sub(net, data_loader=val_loader ,device=\"cuda:0\"):\n    if \"cuda\" in device:\n        net = torch.nn.DataParallel(net).to(device)\n    else:\n        net = net.to(device)\n\n    criterion = nn.CrossEntropyLoss().to(device)\n\n    net.eval()\n    net = net.to(device)\n    losses = AverageMeter()\n    top1 = AverageMeter()\n    top5 = AverageMeter()\n\n    with torch.no_grad():\n        with tqdm(total=len(data_loader), desc=\"Validate\") as t:\n            for i, (images, labels) in enumerate(data_loader):\n                images, labels = images.to(device), labels.to(device)\n                # compute output\n                output = net(images)\n                loss = criterion(output, labels)\n                # measure accuracy and record loss\n                acc1, acc5 = accuracy(output, labels, topk=(1, 5))\n\n                losses.update(loss.item(), images.size(0))\n                top1.update(acc1[0].item(), images.size(0))\n                top5.update(acc5[0].item(), images.size(0))\n                t.set_postfix(\n                    {\n                        \"loss\": losses.avg,\n                        \"top1\": top1.avg,\n                        \"top5\": top5.avg,\n                        \"img_size\": images.size(2),\n                    }\n                )\n                t.update(1)\n\n    print(\n        \"Results: loss=%.5f,\\t top1=%.1f,\\t top5=%.1f\"\n        % (losses.avg, top1.avg, top5.avg)\n    )\n    return top1.avg","metadata":{"_cell_guid":"a67645ff-1f32-4315-91dc-fe16a373e9bc","_uuid":"6c1a6348-670d-4a1a-90f4-f23377d769fc","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-06-01T15:27:18.605008Z","iopub.execute_input":"2024-06-01T15:27:18.605408Z","iopub.status.idle":"2024-06-01T15:27:18.617676Z","shell.execute_reply.started":"2024-06-01T15:27:18.605355Z","shell.execute_reply":"2024-06-01T15:27:18.616197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets replace SE blocks of the network as their current implementation on OFA is not supported on FPGA so they will be replaced by pytorch SE blocks","metadata":{"_cell_guid":"5c3c8c6a-e8c5-4e7a-8bb9-19e6668ed3cb","_uuid":"d6c2941e-c512-4664-91cd-49ef33a0b300","trusted":true}},{"cell_type":"markdown","source":"## 4. Accuracy Predictor \n\nThe key components of very fast neural network deployment are **accuracy predictors** and **efficiency predictors**.\nFor the accuracy predictor, it predicts the Top-1 accuracy of a given sub-network on a **holdout validation set**\n(different from the official 50K validation set) so that we do **NOT** need to run very costly inference on ImageNet\nwhile searching for specialized models. Such an accuracy predictor is trained using an accuracy dataset built with the OFA network.","metadata":{"_cell_guid":"8b6a4d11-ad54-4a39-b51e-0c9659916087","_uuid":"e92941b1-2085-47d7-b7e9-b95357df7005","trusted":true}},{"cell_type":"code","source":"# accuracy predictor\naccuracy_predictor = AccuracyPredictor(\n    pretrained=True,\n    device='cuda:0' if cuda_available else 'cpu'\n)\n\nprint('The accuracy predictor is ready!')\nprint(accuracy_predictor.model)","metadata":{"_cell_guid":"f68fc7af-62e5-46d1-940f-93252791d058","_uuid":"614ed3b7-018f-477c-8618-f87ce0924269","collapsed":false,"jupyter":{"outputs_hidden":false},"pycharm":{"name":"#%%\n"},"execution":{"iopub.status.busy":"2024-06-01T15:27:18.619433Z","iopub.execute_input":"2024-06-01T15:27:18.619882Z","iopub.status.idle":"2024-06-01T15:27:20.317172Z","shell.execute_reply.started":"2024-06-01T15:27:18.619834Z","shell.execute_reply":"2024-06-01T15:27:20.315813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets try predicting accuracy of our randomly sampled subnet  \n\nThis is one network in our search space ","metadata":{"_cell_guid":"73a12c05-0e56-49d8-a7b8-0f0bed3e7efb","_uuid":"5673442d-9c37-48a6-845c-f84e6b2d6718","trusted":true}},{"cell_type":"code","source":"cfg = {'ks': [7, 5, 7, 5, 7, 7, 7, 7, 5, 5, 7, 7, 7, 7, 5, 7, 5, 3, 7, 7],\n   'e': [6, 6, 6, 4, 6, 6, 6, 6, 6, 6, 4, 6, 6, 6, 6, 6, 6, 4, 4, 4],\n   'd': [4, 4, 4, 4, 4]}","metadata":{"execution":{"iopub.status.busy":"2024-06-01T15:27:20.319048Z","iopub.execute_input":"2024-06-01T15:27:20.319546Z","iopub.status.idle":"2024-06-01T15:27:20.326992Z","shell.execute_reply.started":"2024-06-01T15:27:20.31951Z","shell.execute_reply":"2024-06-01T15:27:20.325663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# visualization of network\ndef visualize_subnet(cfg):\n    image_size = 224\n    draw_arch(cfg[\"ks\"], cfg[\"e\"], cfg[\"d\"], image_size, out_name=\"viz/subnet\")\n    im = Image.open(\"viz/subnet.png\")\n    im = im.rotate(90, expand=1)\n    fig = plt.figure(figsize=(im.size[0] / 250, im.size[1] / 250))\n    plt.axis(\"off\")\n    plt.imshow(im)\n    plt.show()\n\n# visualize_subnet(cfg)","metadata":{"execution":{"iopub.status.busy":"2024-06-01T15:27:20.328597Z","iopub.execute_input":"2024-06-01T15:27:20.328992Z","iopub.status.idle":"2024-06-01T15:27:20.343831Z","shell.execute_reply.started":"2024-06-01T15:27:20.328962Z","shell.execute_reply":"2024-06-01T15:27:20.342573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cfg['r']= [224]\n\nacc = accuracy_predictor.predict_accuracy([cfg])\nprint(acc*100)","metadata":{"_cell_guid":"610d4f79-a269-4cf0-bd4b-d09e9e26e7f0","_uuid":"8960373f-1889-4628-93e4-227807439787","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-06-01T15:27:20.34547Z","iopub.execute_input":"2024-06-01T15:27:20.346264Z","iopub.status.idle":"2024-06-01T15:27:20.450706Z","shell.execute_reply.started":"2024-06-01T15:27:20.346227Z","shell.execute_reply":"2024-06-01T15:27:20.449239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cfg","metadata":{"execution":{"iopub.status.busy":"2024-06-01T15:27:20.452259Z","iopub.execute_input":"2024-06-01T15:27:20.452675Z","iopub.status.idle":"2024-06-01T15:27:20.463801Z","shell.execute_reply.started":"2024-06-01T15:27:20.452631Z","shell.execute_reply":"2024-06-01T15:27:20.462319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we have the powerful **accuracy predictor**. We then introduce two types of **efficiency predictors**: the latency predictor and the FLOPs predictor. \n\nThe intuition of having efficiency predictors, especially the latency predictor, is that measuring the latency of a sub-network on-the-fly is also costly, especially for FPGA devices becuase it takes hours of work to implement even a single network and then measure latency.\n\nThe latency predictor is designed to eliminate this cost.","metadata":{"_cell_guid":"8f7f519e-9474-4816-96f6-352e360eec2e","_uuid":"fe970f2f-26c4-4366-80cd-62fc9d5a7814","trusted":true}},{"cell_type":"markdown","source":"## 5. Defining Efficient Arthemetic intensity Constraint\n\nTrying to do three things maximizing the mac operations perfomed by model which represents the learning capacity\nof model while minimizing the number of paramters which is the size of model and also minimizing the latency \non target hardware thus finding a model that has performs maximum calculation while being smaller and faster\n\n","metadata":{"_cell_guid":"84c3dd41-3330-4748-9a91-efe3fc31efef","_uuid":"6b4e342f-e521-4ef6-aaf1-52500decf8db","trusted":true}},{"cell_type":"code","source":"class EfficientCapacityEstimator:\n    \n    def __init__(self, arthemetic_intensity_lookup, latency_estimator):\n        self.ai = arthemetic_intensity_lookup\n        self.lat = latency_estimator\n        \n     # Both latency estmiator and arthemetic intensity calculator are already defined, we will utilize   \n    def predict_efficiency(self, sample):\n        arth_int = 1/self.ai.predict_efficiency(sample)  #actualy returns 1/arth_intensity\n        latency = self.lat.predict_efficiency(sample)\n        \n        efficient_arthemetic_intensity = arth_int / latency\n        \n        # To make the task minimization problem \n        return 1 / efficient_arthemetic_intensity\n","metadata":{"_cell_guid":"62bd0418-2cd8-4b1d-942d-caf9ec792c4b","_uuid":"ce28329d-35e7-4796-9541-7dd09c86b1cb","collapsed":false,"jupyter":{"outputs_hidden":false},"pycharm":{"is_executing":true,"name":"#%%\n"},"execution":{"iopub.status.busy":"2024-06-01T15:27:20.465261Z","iopub.execute_input":"2024-06-01T15:27:20.465759Z","iopub.status.idle":"2024-06-01T15:27:20.477901Z","shell.execute_reply.started":"2024-06-01T15:27:20.465709Z","shell.execute_reply":"2024-06-01T15:27:20.476642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from fpga_utils.latency_estimation import LatencyTable  \n\narthemetic_intensity_lookup = ArthIntTable(pred_type='arthemetic_intensity', \n                                  device='cuda:0' if cuda_available else 'cpu',batch_size=1, \n                                  )\n\nlatency_estimator = LatencyTable()\n\n\nefficiency_estimator = EfficientCapacityEstimator(arthemetic_intensity_lookup, latency_estimator)\n\nprint('The  Efficient Arthemetic intensity predictor is ready!')","metadata":{"execution":{"iopub.status.busy":"2024-06-01T15:27:20.482755Z","iopub.execute_input":"2024-06-01T15:27:20.483226Z","iopub.status.idle":"2024-06-01T15:27:23.103237Z","shell.execute_reply.started":"2024-06-01T15:27:20.48319Z","shell.execute_reply":"2024-06-01T15:27:23.102021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" lets define a custome configuration of network and then try to estimate it's latency and arthemetic intensity","metadata":{}},{"cell_type":"code","source":"print(1/arthemetic_intensity_lookup.predict_efficiency(cfg))\nprint(latency_estimator.predict_efficiency(cfg))\nprint(efficiency_estimator.predict_efficiency(cfg))\n","metadata":{"execution":{"iopub.status.busy":"2024-06-01T15:27:23.104985Z","iopub.execute_input":"2024-06-01T15:27:23.105494Z","iopub.status.idle":"2024-06-01T15:27:23.117028Z","shell.execute_reply.started":"2024-06-01T15:27:23.105451Z","shell.execute_reply":"2024-06-01T15:27:23.115464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Run Evolutionary search\n\nNow, let's proceed towards searching for efficient network under latency constraint. We use the same accuracy predictor since accuracy predictors are agnostic to the types of efficiency constraint . For the efficiency predictor, we build a latency table.","metadata":{"_cell_guid":"2c3dde1c-20cd-41af-b9ce-aeace491a287","_uuid":"7f17cfda-4972-4f1b-8e7a-40b469f0e541","trusted":true}},{"cell_type":"markdown","source":"Lets try these constraint once","metadata":{"_cell_guid":"7e2d5ea2-772d-42f0-863a-61874782a32f","_uuid":"1df4bde8-21a2-4799-892c-80e8d8512e3c","trusted":true}},{"cell_type":"markdown","source":"Note to make the task minimization problem we will search model limiting to contraint 1/ effective_arthemetic_intensity","metadata":{}},{"cell_type":"code","source":"#  Hyper-parameters for the evolutionary search process\n\nP = 100  # The size of population in each generation\nN = 10000  # How many generations of population to be searched\nr = 0.3  # The ratio of networks that are used as parents for next generation\n\nparams = {\n    'constraint_type': 'efficient_arthemetic_intensity', # Efficient Arthemetic intensity constrained search\n    'efficiency_constraint': 0.3,  # suggested range (0.35,4) (min to max), we want smaller and smaller\n    'mutate_prob': 0.3, # The probability of mutation in evolutionary search\n    'mutation_ratio': 0.5, # The ratio of networks that are generated through mutation in generation n >= 2.\n    'efficiency_predictor': efficiency_estimator, # To use a predefined efficiency predictor.\n    'accuracy_predictor': accuracy_predictor, # To use a predefined accuracy_predictor predictor.\n    'population_size': P,\n    'max_time_budget': N,\n    'parent_ratio': r,\n}\n\n\nfinder = EvolutionFinder(**params)\n\n\n# already searched eai_values = [1.4, 1.1, 1, 0.8, 0.65, 0.6]\n#previously searched:[1.4, 1.1, 0.95, 0.8, 0.65]\neai_to_search = [0.55, 0.5, 0.45, 0.4]\n\nresult_lis = []\nresult_valids = []\ninfo = []\nfor efficient_arthemetic_intensity in eai_to_search:\n    \n    print(f\"Starting search for efficient_arthemetic_intensity = {efficient_arthemetic_intensity}\")\n    st = time.time()\n    finder.set_efficiency_constraint(efficient_arthemetic_intensity)\n    best_valids, best_info = finder.run_evolution_search()\n    ed = time.time()\n    \n    print('Found best architecture at efficient_arthemetic_intensity <= %.2f ms*bytes/ops in %.2f seconds! It achieves %.2f%s predicted accuracy with efficient_arthemetic_intensity of %.2f ms*bytes/ops ./n' % (efficient_arthemetic_intensity, ed-st, best_info[0] * 100, '%',best_info[-1]))\n    print(best_info)\n    result_lis.append(best_info)\n    result_valids.append(best_valids)\n    info.append(ed-st)","metadata":{"_cell_guid":"fa0e53fb-4bd0-4030-9e7b-9043a8206a41","_uuid":"337a45a5-262b-4d71-a67b-684f3d4deb79","collapsed":false,"jupyter":{"outputs_hidden":false},"pycharm":{"name":"#%%\n"},"execution":{"iopub.status.busy":"2024-06-01T15:41:42.514688Z","iopub.execute_input":"2024-06-01T15:41:42.515138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(result_lis)","metadata":{"execution":{"iopub.status.busy":"2024-06-01T15:27:56.120166Z","iopub.status.idle":"2024-06-01T15:27:56.120715Z","shell.execute_reply.started":"2024-06-01T15:27:56.12047Z","shell.execute_reply":"2024-06-01T15:27:56.120492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The results after runnint them are !\n\nLets visualize and evalute the found architecutre","metadata":{"_cell_guid":"a123c81b-3a27-4aa5-9e57-98960e64985f","_uuid":"915a9451-60c8-4a1d-a5d4-a6157e47b2a6","trusted":true}},{"cell_type":"code","source":"for arch in result_lis:\n    _,cfg, _ = models_cfg[arch]\n    print(cfg)\n    visualize_subnet(cfg)\n    \n    latency = latency_estimator.predict_efficiency(cfg)\n    ai = 1 / arthemetic_intensity_lookup.predict_efficiency(cfg)\n    accuracy = accuracy_predictor.predict_accuracy([cfg]).item() * 100\n    print(f\"Latency: {latency:.2f}ms, AI: {ai:.2f}ops/byte, Accuracy: {accuracy:.2f}%, CLI: {cli:.3f}\")\n    ","metadata":{"execution":{"iopub.status.busy":"2024-02-14T04:27:34.627747Z","iopub.status.idle":"2024-02-14T04:27:34.628104Z","shell.execute_reply.started":"2024-02-14T04:27:34.627936Z","shell.execute_reply":"2024-02-14T04:27:34.627952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Previous Search results architectures:\nThese are the architectures we obtained from previous search with following constraint.  \n\nTooks hour of searching, Nearly 25gpu hours for all these networks ","metadata":{}},{"cell_type":"code","source":"eai_values = [1.4, 1.1, 1, 0.8, 0.65, 0.6]\n\nmodels_cfg = {\n\"1.4\": (0.8506609201431274, {'wid': None, 'ks': [5, 5, 5, 5, 5, 5, 7, 7, 5, 7, 3, 5, 7, 7, 5, 7, 5, 3, 5, 5], 'e': [4, 6, 6, 6, 6, 6, 6, 6, 4, 3, 4, 6, 6, 6, 6, 6, 6, 6, 6, 4], 'd': [4, 4, 4, 4, 4], 'r': [224]}, 1.2025607975380936),   \n\"1.1\": (0.8530834913253784, {'wid': None, 'ks': [5, 5, 3, 5, 5, 3, 7, 3, 7, 7, 3, 7, 5, 7, 5, 3, 5, 3, 5, 5], 'e': [4, 6, 4, 4, 6, 6, 4, 6, 4, 3, 6, 6, 6, 6, 6, 3, 6, 6, 6, 3], 'd': [4, 4, 4, 4, 4], 'r': [224]}, 1.0759729786277912),   \n\"1.0\":(0.8514131307601929, {'wid': None, 'ks': [5, 5, 3, 5, 5, 3, 7, 3, 7, 7, 3, 7, 7, 7, 3, 3, 5, 3, 5, 5], 'e': [4, 3, 4, 4, 6, 6, 4, 6, 4, 3, 6, 6, 6, 6, 6, 4, 6, 6, 3, 3], 'd': [4, 4, 4, 4, 4], 'r': [224]}, 0.9969672772703635),    \n\"0.8\": (0.8426103591918945, {'wid': None, 'ks': [3, 3, 3, 5, 3, 3, 3, 5, 7, 3, 5, 3, 3, 5, 3, 3, 5, 3, 5, 3], 'e': [4, 3, 3, 3, 6, 6, 4, 4, 4, 6, 6, 4, 6, 6, 6, 6, 6, 4, 4, 6], 'd': [4, 4, 4, 4, 3], 'r': [224]}, 0.7998816261063237),\n\"0.65\": (0.8276481628417969, {'wid': None, 'ks': [3, 3, 3, 3, 3, 3, 5, 3, 7, 3, 3, 3, 3, 5, 3, 3, 7, 3, 5, 5], 'e': [4, 6, 6, 4, 6, 6, 4, 4, 4, 4, 4, 3, 6, 6, 3, 6, 3, 4, 4, 3], 'd': [3, 4, 4, 4, 2], 'r': [224]}, 0.6497582036831724),\n\"0.6\":  (0.822089433670044, {'wid': None, 'ks': [3, 3, 3, 7, 3, 3, 3, 3, 5, 3, 5, 5, 3, 3, 3, 3, 3, 3, 5, 7], 'e': [4, 6, 3, 3, 6, 4, 4, 4, 4, 3, 4, 3, 4, 4, 6, 3, 4, 3, 3, 4], 'd': [3, 4, 3, 4, 2], 'r': [224]}, 0.599851266295482)\n    \n}","metadata":{"execution":{"iopub.status.busy":"2024-06-01T06:00:11.554432Z","iopub.execute_input":"2024-06-01T06:00:11.554978Z","iopub.status.idle":"2024-06-01T06:00:11.579237Z","shell.execute_reply.started":"2024-06-01T06:00:11.554913Z","shell.execute_reply":"2024-06-01T06:00:11.577672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for arch in models_cfg:\n    _,cfg, cli = models_cfg[arch]\n    visualize_subnet(cfg)\n    latency = latency_estimator.predict_efficiency(cfg)\n    ai = 1 / arthemetic_intensity_lookup.predict_efficiency(cfg)\n    accuracy = accuracy_predictor.predict_accuracy([cfg]).item() * 100\n    print(f\"Latency: {latency:.2f}ms, AI: {ai:.2f}ops/byte, Accuracy: {accuracy:.2f}%, CLI: {cli:.3f}\")","metadata":{"execution":{"iopub.status.busy":"2024-06-01T06:00:13.724902Z","iopub.execute_input":"2024-06-01T06:00:13.725434Z","iopub.status.idle":"2024-06-01T06:00:17.194609Z","shell.execute_reply.started":"2024-06-01T06:00:13.72539Z","shell.execute_reply":"2024-06-01T06:00:17.193443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!mkdir Models","metadata":{"_cell_guid":"d76ade3e-8656-4e04-9a49-fca08df3b51c","_uuid":"1374eab8-6211-445d-b5ab-f44af98757b7","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.632812Z","iopub.status.idle":"2024-02-14T04:27:34.633121Z","shell.execute_reply.started":"2024-02-14T04:27:34.632968Z","shell.execute_reply":"2024-02-14T04:27:34.632981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7. Visualizing and Evaluating Searched Models\nLets build model from config and Replace SE blocks in all these networks","metadata":{"_cell_guid":"5c8cfbf1-5cb8-498f-8110-b6efa3fa7f92","_uuid":"a89eb086-999e-4f0b-83d3-e206a80f0da8","trusted":true}},{"cell_type":"markdown","source":"We will be utilizing **[Once-for-All (OFA)](https://github.com/mit-han-lab/once-for-all)** network trained on MobilenetV3 search space as a our supernetwork to intialize the weights which would make the training of our searched network faster.","metadata":{}},{"cell_type":"code","source":"%%script echo skipping\nnet_id  = 'ofa_mbv3_d234_e346_k357_w1.2'\nurl_base = \"https://raw.githubusercontent.com/han-cai/files/master/ofa/ofa_nets/\"\n\nofa_network = OFAMobileNetV3(\n            dropout_rate=0,\n            width_mult=1.2,\n            ks_list=[3, 5, 7],\n            expand_ratio_list=[3, 4, 6],\n            depth_list=[2, 3, 4],\n        )\n\npt_path = download_url(url_base + net_id, model_dir=\".torch/ofa_nets\")\ninit = torch.load(pt_path, map_location=\"cpu\")[\"state_dict\"]\nofa_network.load_state_dict(init)\nprint('Supernetwork Ready')","metadata":{"execution":{"iopub.status.busy":"2024-02-14T04:27:34.635115Z","iopub.status.idle":"2024-02-14T04:27:34.63546Z","shell.execute_reply.started":"2024-02-14T04:27:34.635288Z","shell.execute_reply":"2024-02-14T04:27:34.635302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script echo skipping\n\nprint('Replacing SE blocks of OFA and \\n Saving the searched models ')\n\nfor arch in result_lis:\n    cfg = arch[1]\n    visualize_subnet(cfg)  \n    ofa_network.set_active_subnet(cfg['ks'], cfg['e'], cfg['d'])\n    network = ofa_network.get_active_subnet()\n\n    print('--'*50)\n    \nmodels = {}   # eai, network pair\nfor i, arch in enumerate(result_lis):\n#     cfg = arch[1]\n    cfg = arch\n    ofa_network.set_active_subnet(cfg['ks'], cfg['e'], cfg['d'])\n    network = ofa_network.get_active_subnet(preserve_weight=True)\n    models[eai_list[i]] = network\n\nfor key, model in models.items():\n    replace_all(model)\n    name = f\"Models/model_search_{key}.pth\"\n    torch.save(model, name)\n    print('Done')","metadata":{"_cell_guid":"b320c23f-93d8-4ee3-93bd-a1765f279b2a","_uuid":"3926e24e-3ff7-469f-8476-4f529d433cfa","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.63651Z","iopub.status.idle":"2024-02-14T04:27:34.636873Z","shell.execute_reply.started":"2024-02-14T04:27:34.636707Z","shell.execute_reply":"2024-02-14T04:27:34.636722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Saving architecture configuration to a File**","metadata":{"_cell_guid":"6900c50d-505a-4725-9198-7a556049821b","_uuid":"6b55b809-8ea2-4f07-9f6a-7fec3def684b","trusted":true}},{"cell_type":"code","source":"import json\n\ndata = {\n    \"arch\": result_lis,\n    \"info\": info\n}\n\nwith open('search_data.json', 'w') as file:\n    json.dump(data, file)\n    \nprint('The cofig data is exported, \\n The exported Data is:')\nprint(data)","metadata":{"_cell_guid":"3c5e3db6-6649-4620-b395-103904eec52b","_uuid":"b3c4e670-fd9a-4101-9026-857c8eaa6ddc","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.638404Z","iopub.status.idle":"2024-02-14T04:27:34.638753Z","shell.execute_reply.started":"2024-02-14T04:27:34.63856Z","shell.execute_reply":"2024-02-14T04:27:34.638573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8. Retrain the Searched network\nLets fine tune the network on a subset of training set of imagenet.","metadata":{"_cell_guid":"afbbdb50-bf35-4170-97fd-fa2f2355b0c0","_uuid":"8e200e47-208e-4c1a-9284-c02e0ed70a0b","trusted":true}},{"cell_type":"markdown","source":"I previously performed FPGA aware Neural architecutral search and found different models with different latency through evolutionary algorithm.\n\nThe obtained network from search process will be now retrained to improve accuracy. Also these models wt were intialized with the help of OFA!","metadata":{"_cell_guid":"82bd3e62-0f92-4d26-9991-ce9c8d26ce2b","_uuid":"64743ae8-1816-49bc-8d80-b261d4423cfb","trusted":true}},{"cell_type":"code","source":"%%script echo skipping\nimport torch \nimport torchvision\nimport os\nimport torch.nn as nn\nfrom torchvision import transforms, datasets\nimport math\nimport time\nfrom tqdm import tqdm\nimport shutil\nimport matplotlib.pyplot as plt","metadata":{"_cell_guid":"a2e3a91a-8119-496c-85c6-df09ac11a018","_uuid":"33432336-296b-464f-bd14-6e39ea995376","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.640009Z","iopub.status.idle":"2024-02-14T04:27:34.640347Z","shell.execute_reply.started":"2024-02-14T04:27:34.640177Z","shell.execute_reply":"2024-02-14T04:27:34.640191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script echo skipping\ncuda_available = torch.cuda.is_available()\nif cuda_available:\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.benchmark = True\n    print('Using GPU.')\nelse:\n    print('Using CPU.')","metadata":{"_cell_guid":"995ad112-0da3-453d-981c-7ffb3dac67d0","_uuid":"9f60a22d-8932-4b20-909f-0f72e11d61a2","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.641278Z","iopub.status.idle":"2024-02-14T04:27:34.641635Z","shell.execute_reply.started":"2024-02-14T04:27:34.641446Z","shell.execute_reply":"2024-02-14T04:27:34.641461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script echo skipping\nbatch_size=32\n\n#I will use a susbset of imagenetval of 10k images \nif cuda_available:\n    # path to the ImageNet dataset\n    # link --> https://www.kaggle.com/datasets/titericz/imagenet1k-val\n    \n    imagenet_data_path = '/kaggle/input/imagenet1k-subset-100k-train-and-10k-val/imagenet_subval'\n\n    # if 'imagenet_data_path' is empty, download a subset of ImageNet containing 2000 images (~250M) for test\n    if not os.path.isdir(imagenet_data_path):\n        print('%s is empty. Download a subset of ImageNet for test.' % imagenet_data_path)\n\n    print('The ImageNet dataset files are ready.')\nelse:\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')\n    \n    \n  \nif cuda_available:\n    # The following function build the data transforms for test\n    def build_val_transform(size):\n        return transforms.Compose([\n            transforms.Resize(int(math.ceil(size / 0.875))),\n            transforms.CenterCrop(size),\n            transforms.ToTensor(),\n            transforms.Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225]\n            ),\n        ])\n    \n    val_data = datasets.ImageFolder(\n            root=os.path.join(imagenet_data_path),\n            transform=build_val_transform(224)\n        )\n    \n\n    val_loader = torch.utils.data.DataLoader(\n        val_data,\n        batch_size=batch_size,  \n        shuffle = True,\n        num_workers=4,  \n        pin_memory=True,\n        drop_last=False,\n    )\n    print('The ImageNet dataloader is ready. Size : {}'.format(len(val_loader)*batch_size))\nelse:\n    data_loader = None\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')","metadata":{"_cell_guid":"b49ecee7-e265-4267-9cb1-020ef5a8d2cb","_uuid":"a6cdacee-d402-4b3e-9ac4-fa080dd481c0","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.643549Z","iopub.status.idle":"2024-02-14T04:27:34.644007Z","shell.execute_reply.started":"2024-02-14T04:27:34.643778Z","shell.execute_reply":"2024-02-14T04:27:34.643797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script echo skipping\ntrain_path = '/kaggle/input/imagenet1k-subset-100k-train-and-10k-val/imagenet_subtrain'\n\ntrain_data = datasets.ImageFolder(\n            root= train_path,\n            transform=build_val_transform(224)\n        )\n\ntrain_loader = torch.utils.data.DataLoader(\n        train_data,\n        batch_size=batch_size, \n        shuffle = True,\n        num_workers=4,  \n        pin_memory=True,\n        drop_last=False,\n    )\n\nprint('The ImageNet train set is ready. Size : {}'.format(len(train_loader)*batch_size))","metadata":{"_cell_guid":"a6f8888b-b40a-4741-9781-62dd5ca523d0","_uuid":"36496cc1-9614-4f63-81d1-ee99e583cc2e","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.645269Z","iopub.status.idle":"2024-02-14T04:27:34.645743Z","shell.execute_reply.started":"2024-02-14T04:27:34.64549Z","shell.execute_reply":"2024-02-14T04:27:34.645508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script echo skipping\ndataloaders = {}\ndataloaders['train'] = train_loader\ndataloaders['val'] = val_loader\n\ndataset_sizes = {'train': len(train_loader)*32,\n                'val': len(val_loader)*32}\nprint(dataset_sizes)","metadata":{"_cell_guid":"b768509f-197f-4991-8013-413fbaedaf22","_uuid":"e3d57b25-eae9-4430-a9c7-9f49c8dd55ad","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.647096Z","iopub.status.idle":"2024-02-14T04:27:34.647546Z","shell.execute_reply.started":"2024-02-14T04:27:34.647306Z","shell.execute_reply":"2024-02-14T04:27:34.647326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n%%script echo skipping\ndef train_model(model, criterion, optimizer, scheduler, num_epochs=5):\n    since = time.time()\n\n    #storing epoch data\n    epoch_data = {\n        'epoch': [],\n        'train': {'loss': [], 'top1_acc': [], 'top5_acc': []},\n        'val': {'loss': [], 'top1_acc': [], 'top5_acc': []}\n    }\n    \n    # Create a temporary directory\n    tempdir = '/kaggle/working/temp'\n    os.makedirs(tempdir, exist_ok=True)\n    best_model_params_path = os.path.join(tempdir, 'best_model_params.pt')\n\n    torch.save(model.state_dict(), best_model_params_path)\n    best_acc = 0.0\n\n    for epoch in range(num_epochs):\n        print(f'Epoch {epoch+1}/{num_epochs}')\n        print('-' * 10)\n        epoch_data['epoch'].append(epoch+1)\n        \n        for phase in ['train', 'val']:\n            if phase == 'train':\n                model.train()\n            else:\n                model.eval()\n            running_loss = 0.0\n            top1_corrects = 0\n            top5_corrects = 0\n            \n            for inputs, labels in tqdm(dataloaders[phase], leave=False):\n                inputs = inputs.to(device)\n                labels = labels.to(device)\n\n                optimizer.zero_grad()\n\n                with torch.set_grad_enabled(phase == 'train'):\n                    outputs = model(inputs)\n                    _, preds = torch.max(outputs, 1)\n                    loss = criterion(outputs, labels)\n\n                    if phase == 'train':\n                        loss.backward()\n                        optimizer.step()\n\n                running_loss += loss.item() * inputs.size(0)\n                \n                # Calculate top-1 accuracy\n                top1_corrects += torch.sum(preds == labels.data)\n                \n                # Calculate top-5 accuracy\n                _, top5_preds = torch.topk(outputs, 5, dim=1)\n                top5_corrects += torch.sum(top5_preds == labels.view(-1, 1))\n\n            if phase == 'train':\n                scheduler.step()\n\n            epoch_loss = running_loss / dataset_sizes[phase]\n            epoch_top1_acc = top1_corrects.double() / dataset_sizes[phase]\n            epoch_top5_acc = top5_corrects.double() / dataset_sizes[phase]\n            \n            epoch_data[phase]['loss'].append(epoch_loss)\n            epoch_data[phase]['top1_acc'].append(epoch_top1_acc)\n            epoch_data[phase]['top5_acc'].append(epoch_top5_acc)\n\n            print(f'{phase} Loss: {epoch_loss:.4f} Top-1 Acc: {epoch_top1_acc:.4f} Top-5 Acc: {epoch_top5_acc:.4f}')\n\n            if phase == 'val' and epoch_top1_acc > best_acc:\n                best_acc = epoch_top1_acc\n                best_top5 = epoch_top5_acc\n                torch.save(model.state_dict(), best_model_params_path)\n\n        print()\n\n    time_elapsed = time.time() - since\n    print(f'Training complete in {time_elapsed // 60:.0f}m {time_elapsed % 60:.0f}s')\n    print(f'Best val Top-1 Acc {best_acc:4f} /n Best val Top-5: {best_top5:4f}')\n\n    model.load_state_dict(torch.load(best_model_params_path))\n\n    # Clean up the temporary directory\n    shutil.rmtree(tempdir)\n\n    return model, epoch_data\n","metadata":{"_cell_guid":"9c1f3261-c421-412c-832e-0430781f8477","_uuid":"ec1f38ea-c450-466b-b68f-8b69452a9434","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.64968Z","iopub.status.idle":"2024-02-14T04:27:34.650023Z","shell.execute_reply.started":"2024-02-14T04:27:34.64986Z","shell.execute_reply":"2024-02-14T04:27:34.649874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%script echo skipping\n\ntorch.cuda.empty_cache()\n# model = torch.load('Models/model_search_11.pth')\nmodel = network.to(device)\n\ncriterion = nn.CrossEntropyLoss()\n# Observe that all parameters are being optimized\noptimizer_ft = torch.optim.SGD(model.parameters(), lr=0.00001, momentum=0.9)\n\n# Decay LR by a factor of 0.1 every 7 epochs\nexp_lr_scheduler = torch.optim.lr_scheduler.StepLR(optimizer_ft, step_size=2, gamma=0.5)\n\nmodel, epoch_data = train_model(model, criterion, optimizer_ft, exp_lr_scheduler,\n                   num_epochs=2)\n\ntorch.save(model, 'retrained_model_search_11.pth')\nprint('*****************************************************')\nprint(epoch_data)","metadata":{"execution":{"iopub.status.busy":"2024-02-14T04:27:34.650984Z","iopub.status.idle":"2024-02-14T04:27:34.651281Z","shell.execute_reply.started":"2024-02-14T04:27:34.65113Z","shell.execute_reply":"2024-02-14T04:27:34.651143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### It is clear that accuracy is recoverable and we can obtain comparable accuracy with our efficiency contraints by retraining for few epochs","metadata":{"_cell_guid":"ff3894d4-cec7-4720-b42f-d6d594dcc950","_uuid":"1e7aec39-a70e-4250-b393-e5c5e2476e8b","trusted":true}},{"cell_type":"markdown","source":"## 6. Insights and Comparision","metadata":{"_cell_guid":"50055b45-a44f-4f89-8712-69e0f9c96819","_uuid":"739e1062-7782-450e-961b-e33441bd4165","trusted":true}},{"cell_type":"code","source":"# 'ResNet-152':81.3,\n#  'Efficientnet V2 Small': 83.6,","metadata":{"_cell_guid":"1c1735e3-d8f1-4274-9e6c-6ffc2744c0b7","_uuid":"62cdea74-7374-4e1c-ad22-2f36c4ee1197","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.652845Z","iopub.status.idle":"2024-02-14T04:27:34.653172Z","shell.execute_reply.started":"2024-02-14T04:27:34.653013Z","shell.execute_reply":"2024-02-14T04:27:34.653027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import shutil\n\n# def zip_folder(folder_path, zip_path):\n#     try:\n#         # Zip the entire folder and its contents\n#         shutil.make_archive(zip_path, 'zip', folder_path)\n#         print(f'Folder \"{folder_path}\" successfully zipped to \"{zip_path}.zip\"')\n#     except Exception as e:\n#         print(f'Error zipping folder: {e}')\n\n# # Example usage:\n\n# zip_folder('/kaggle/working/Models/', '/kaggle/working/models')","metadata":{"_cell_guid":"7e3d79eb-eb5c-4365-b914-9b60a52247f6","_uuid":"4d002645-9da2-4470-8896-ddbc092a283a","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.654458Z","iopub.status.idle":"2024-02-14T04:27:34.655057Z","shell.execute_reply.started":"2024-02-14T04:27:34.65481Z","shell.execute_reply":"2024-02-14T04:27:34.654831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# archs = [\n# ['Efficientnet B0', 77.8, 19.5],\n# ['Mobilenet V3_large', 77, 11.5],\n# ['nvidia_efficientnet_b4',82, 20.3]\n# ]\n\n\n# y = [res[0]*100 for res in result_lis]+ [79.5] #added evaulated accuracy after one epcoh train\n# x = [1/res[2] for res in result_lis]+ [20.1]\n# plt.figure(figsize=(12, 8))\n# plt.xlabel('Arithmetic Intensity (ops/byte)', fontsize=14)\n# plt.ylabel('Predicted Accuracy', fontsize=14)\n\n# # Plot the line and scatter plot with 'X' marker for each point in the same color\n# plt.plot(x, y, marker='X', linestyle='-', color='blue', markersize=8, label = 'Ours')\n# plt.plot(archs[0][2], archs[0][1], marker='*', color='red', label ='other', markersize=10 )\n# plt.plot(archs[1][2],archs[1][1], marker = '*', color = 'red', markersize = 10)\n# plt.plot(archs[2][2],archs[2][1], marker = '*', color = 'red', markersize = 10)\n\n\n# # Annotate the middle point\n# plt.text(20.2,79, f' Evaluated Accuracy \\n after 1 epoch', fontsize = 10)\n# plt.text(17, 85.2, f'      FPGA specialized networks', fontsize=12, verticalalignment='bottom', horizontalalignment='left', color='blue')\n# plt.text(archs[0][2], archs[0][1], f'  Efficientnet_b0', fontsize=12 )\n# plt.text(archs[1][2],archs[1][1], f' {archs[1][0]}')\n# plt.text(archs[2][2]+0.1,archs[2][1], f' {archs[2][0]}', fontsize=12)\n\n# plt.ylim(75, 89)\n# plt.xlim(10, 24)\n# # Add legend\n# plt.legend(fontsize=12)\n# plt.title('Searching for FPGAs')\n# # Show the plot\n# plt.show()","metadata":{"_cell_guid":"6a2e9b32-330b-4478-b0d9-82f6c74b763e","_uuid":"3264d105-89ed-4d3a-9ee1-51f3e4a0c946","collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2024-02-14T04:27:34.656116Z","iopub.status.idle":"2024-02-14T04:27:34.65656Z","shell.execute_reply.started":"2024-02-14T04:27:34.656325Z","shell.execute_reply":"2024-02-14T04:27:34.656344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_cell_guid":"8d6227c6-7a8f-4332-a53b-109b39a118f0","_uuid":"fe3c0f71-9689-49e4-a435-2bae691a423c","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]}]}