{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":6799,"databundleVersionId":4225553,"sourceType":"competition"},{"sourceId":7556785,"sourceType":"datasetVersion","datasetId":4400759},{"sourceId":7582952,"sourceType":"datasetVersion","datasetId":4413931}],"dockerImageVersionId":30648,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Searching for Hardware Aware Neural Networks on ImageNet for FPGA in MobilenetV3 Search Space\n\n\nWe will be utilizing **[Once-for-All (OFA)](https://github.com/mit-han-lab/once-for-all)** network trained on MobilenetV3 search space as a our supernetwork to intialize the weights which would make the training of our searched network faster. We also need to estimate the accuracy of samples in our search space because it requires significant amount of time to train each subnet for evaluating its perfomace. This slows down the searching process. Utilizing the accuracy peredictory MLP that was trained on the (arch, accu) dataset in our search space would speed up the search process as we won't have to train the network at all. The search is perfomed on the Imagenet Dataset. \n\nWe have also build a latency table by manualy deploying all the unique blocks of our search space on ULTRA96v2 FPGA board this will be utilized to estimate latency of our networks.","metadata":{"_uuid":"214f39c3-e9b3-4d4a-9191-7fbccd0da632","_cell_guid":"b72d2cc2-6903-4a1a-b966-dd65184120a5","trusted":true}},{"cell_type":"markdown","source":"## 1. Preparation\nLet's first install all the required packages:","metadata":{"_uuid":"ca674990-a38b-40ff-9e89-cc0242feac5a","_cell_guid":"7ff1f4fe-c4e8-464e-bcd0-f05d4a2fb091","trusted":true}},{"cell_type":"code","source":"# For kaggle\n!rm -r /kaggle/working/Evolutionary-Neural-Architectural-Search-for-FPGAs /kaggle/working/ofa /kaggle/working/viz /kaggle/working/search_space_blocks /kaggle/working/blocks /kaggle/working/figures\n\n!pip install thop \n! pip install gdown\n!pip install shutil\n!pip install graphviz\n! pip install torch-summary \n! git clone --branch for_kaggle https://github.com/amitpant7/Evolutionary-Neural-Architectural-Search-for-FPGAs.git\n! mv -f /kaggle/working/Evolutionary-Neural-Architectural-Search-for-FPGAs/* /kaggle/working\n! rm -r Evolutionary-Neural-Architectural-Search-for-FPGAs","metadata":{"_uuid":"2541a397-5182-423f-a65e-a83893b1a12a","_cell_guid":"5f7d2c59-4621-45ce-bf03-40e532baaf47","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install thop \n! pip install gdown\n!pip install shutil\n!pip install graphviz\n! pip install torch-summary","metadata":{"_uuid":"e0f9376e-390f-47e2-8026-06088f0c7457","_cell_guid":"e8f24185-24c8-4f69-98fc-2234a3274353","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Installing PyTorch...')\n! pip install torch \nprint('Installing torchvision...')\n! pip install torchvision\nprint('Installing numpy...')\n! pip install numpy \n# thop is a package for FLOPs computing.\nprint('Installing thop (FLOPs counter) ...')\n! pip install thop \n# ofa is a package containing training code, pretrained specialized models and inference code for the once-for-all networks.\n# print('Installing OFA...')\n# ! pip install ofa \n# tqdm is a package for displaying a progress bar.\nprint('Installing tqdm (progress bar) ...')\n! pip install tqdm \nprint('Installing matplotlib...')\n! pip install matplotlib \n! pip install torch-summary \n\nprint('All required packages have been successfully installed!')","metadata":{"_uuid":"c79b023d-9eb3-4fd9-87d4-624f89997d7d","_cell_guid":"51e2d4b2-f708-426b-8006-f93ac6f1de51","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Then, we can import the packages used in this tutorial:","metadata":{"_uuid":"bd7c56c3-f0d5-4cc1-886f-80adb0b34918","_cell_guid":"f552a9b7-c47b-4998-9e7a-376e85cc8737","trusted":true}},{"cell_type":"code","source":"import os\nimport torch\nimport torch.nn as nn\nfrom torchvision import transforms, datasets\nimport numpy as np\nimport time\nimport random\nimport shutil\nimport math\nfrom PIL import Image\nimport copy\nfrom matplotlib import pyplot as plt\nfrom torchsummary import summary\n\nfrom ofa.model_zoo import ofa_net\nfrom ofa.utils import download_url\n\nfrom ofa.accuracy_predictor import AccuracyPredictor\nfrom ofa.flops_table import ArthIntTable\n\nfrom ofa.latency_table import LatencyTable\nfrom ofa.evolution_finder import EvolutionFinder\nfrom ofa.imagenet_eval_helper import evaluate_ofa_subnet, evaluate_ofa_specialized\nfrom ofa.imagenet_classification.elastic_nn.networks.ofa_mbv3 import OFAMobileNetV3\n\nfrom ofa.utils.arch_visualization_helper import draw_arch\n# from ofa.tutorial import AccuracyPredictor, FLOPsTable, LatencyTable, EvolutionFinder\n# from ofa.tutorial import evaluate_ofa_subnet, evaluate_ofa_specialized\n\nfrom tqdm import tqdm\n\n# set random seed\nrandom_seed = 1\nrandom.seed(random_seed)\nnp.random.seed(random_seed)\ntorch.manual_seed(random_seed)\nprint('Successfully imported all packages and configured random seed to %d!'%random_seed)","metadata":{"_uuid":"ef7c9af0-0bcc-4cd2-827d-a2e219c22dc4","_cell_guid":"ed77b74d-713b-40d2-9248-96d4a9556ad4","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{"_uuid":"86973de5-4c72-492e-adf6-17333c7071ac","_cell_guid":"f279ef50-a6ab-4a80-80ce-709e1ec8cffb","trusted":true}},{"cell_type":"markdown","source":"Now it's time to determine which device to use for neural network inference in the rest of this tutorial. If your machine is equipped with GPU(s), we will use the GPU by default. Otherwise, we will use the CPU.","metadata":{"_uuid":"6cfbc39e-7dfd-4258-85a0-2b33f71a1aae","_cell_guid":"0848ed15-078c-4fbb-90e8-0845550faec4","trusted":true}},{"cell_type":"code","source":"#os.environ['CUDA_VISIBLE_DEVICES'] = '0'\ncuda_available = torch.cuda.is_available()\nif cuda_available:\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.benchmark = True\n    torch.cuda.manual_seed(random_seed)\n    print('Using GPU.')\nelse:\n    print('Using CPU.')","metadata":{"_uuid":"dbf7d111-1be6-4b50-9e06-352221ee2dbc","_cell_guid":"d5958bd8-92cb-47f5-881b-aa17cf37efd3","collapsed":false,"pycharm":{"name":"#%%\n"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  2. Architecutre Visualization & Encoding: Exploring the OFA network","metadata":{"_uuid":"c09e05ba-fe8f-40f4-9e5a-a1571e6f49cb","_cell_guid":"bb6c6dd9-f5f0-4827-9910-960255b1adc6","trusted":true}},{"cell_type":"markdown","source":"Good! Now you have successfully configured the environment! It's time to import the **OFA network** for the following experiments.\nThe OFA network used in this tutorial is built upon MobileNetV3 with width multiplier 1.2, supporting elastic depth (2, 3, 4) per stage, elastic expand ratio (3, 4, 6), and elastic kernel size (3, 5 7) per block.","metadata":{"_uuid":"89b99fd6-0852-448d-8320-3ad97db129d8","_cell_guid":"4e9a897b-2400-4dfb-b645-4c6e7d50b283","trusted":true}},{"cell_type":"code","source":"net_id  = 'ofa_mbv3_d234_e346_k357_w1.2'\nurl_base = \"https://raw.githubusercontent.com/han-cai/files/master/ofa/ofa_nets/\"\n\nofa_network = OFAMobileNetV3(\n            dropout_rate=0,\n            width_mult=1.2,\n            ks_list=[3, 5, 7],\n            expand_ratio_list=[3, 4, 6],\n            depth_list=[2, 3, 4],\n        )\n\npt_path = download_url(url_base + net_id, model_dir=\".torch/ofa_nets\")\ninit = torch.load(pt_path, map_location=\"cpu\")[\"state_dict\"]\nofa_network.load_state_dict(init)\nprint('Supernetwork Ready')","metadata":{"_uuid":"56ca4498-3d2c-4979-b432-676d4ea3d4f4","_cell_guid":"671651e1-28c6-465b-b642-26aee18857b2","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets visualize a randomly sampled network from our supernet.\n\nIn the architecture visualization, the legend of each block MBConv{e}-{k}x{k} means that the current block is a mobile inverted block with expand ratio e and the kernel size of the depthwise convolution layer is k. Different colors of the blocks indicate different kernel sizes, and gray blocks are network stage dividers. Different widths for the blocks indicate different expand ratios. We also annotate the output resolution close to each block.","metadata":{"_uuid":"7f5ba9b9-d1a6-49dd-a433-7a7ee8790783","_cell_guid":"bab6acd0-f2a9-455b-a6c1-57cdcc5cdf1a","trusted":true}},{"cell_type":"code","source":"# Randomly sample sub-networks from OFA network\nimage_size = 224\n\ncfg1 = ofa_network.sample_active_subnet()\nsubnet = ofa_network.get_active_subnet(preserve_weight=True)\n\n#manualy set the subnet \ncfg = ofa_network.set_active_subnet(ks=3, e=6, d=4)\n\ncfg = ofa_network.set_max_net()\nsubnet2 = ofa_network.get_active_subnet(preserve_weight=True)","metadata":{"_uuid":"46f3910d-a8b9-48d1-890f-c78d4cda4afc","_cell_guid":"3c96ea44-9ab4-4bb1-a6c1-8df5db5e2d2f","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def visualize_subnet(cfg):\n    draw_arch(cfg[\"ks\"], cfg[\"e\"], cfg[\"d\"], image_size, out_name=\"viz/subnet\")\n    im = Image.open(\"viz/subnet.png\")\n    im = im.rotate(90, expand=1)\n    fig = plt.figure(figsize=(im.size[0] / 250, im.size[1] / 250))\n    plt.axis(\"off\")\n    plt.imshow(im)\n    plt.show()\n\nvisualize_subnet(cfg)","metadata":{"_uuid":"d70fb607-b2e1-445c-9934-f550a27e3244","_cell_guid":"3bdf9a03-0503-449b-868b-88009805638b","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" \nEvery subnet in the ofa can be represented in the form of python dictionary. This encoding helps us to represent entire netowk with just few numbers.\n\nHere is one expample of encoding","metadata":{"_uuid":"109b284d-ac1f-45f0-954d-15c94d5a5c1c","_cell_guid":"e9e4f636-0f27-42f6-a86d-bce139e98ee6","trusted":true}},{"cell_type":"code","source":"print('The architecture encoding of random subnetwork', cfg)","metadata":{"_uuid":"e9494c90-f08b-4cae-8207-6c35e0210531","_cell_guid":"427543e2-c7a9-485f-99bc-1be0bf4c3cdb","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Dataset Preperation","metadata":{"_uuid":"4db057a3-8141-4ee3-98d7-675463897305","_cell_guid":"8bbdc985-a6ba-4983-9df7-2105d9ba771b","trusted":true}},{"cell_type":"markdown","source":"Now, let's build the ImageNet dataset and the corresponding dataloader. Notice that **if you are not in kaggle it will be skipped** since it will be very slow.\n\n\nWe will only use subset of ImageNet validation set which will contains 10,000 images for testing.","metadata":{"_uuid":"07c22216-4e65-4f25-9bd2-b1ac2f6d0e57","_cell_guid":"ebaa20c6-25ae-4f12-9982-1e7fdadee3bc","pycharm":{"name":"#%% md\n"},"trusted":true}},{"cell_type":"markdown","source":"**I will utilize subsets of imagenet for both validation and retraining.**","metadata":{"_uuid":"919942ef-4c54-48ef-9f1f-aedcdaa044df","_cell_guid":"d5ff6436-4b58-45e7-8c39-1a4c00b171d3","trusted":true}},{"cell_type":"code","source":"# #creating subset of 100k using imagenet original.\n# import os\n# import shutil\n# import random\n# from multiprocessing import Pool\n\n# def copy_files(src_file_path, tgt_file_path):\n#     try:\n#         shutil.copy(src_file_path, tgt_file_path)\n#     except Exception as e:\n#         print(f\"Error copying {src_file_path} to {tgt_file_path}: {e}\")\n\n\n# def make_subset(old_path, new_path, number_per_class=10, num_processes=4, batch_size=256):\n#     if os.path.exists(new_path):\n#         shutil.rmtree(new_path)\n\n#     os.makedirs(new_path)\n\n#     print('Creating subset...')\n#     dirs = [d for d in os.listdir(old_path) if os.path.isdir(os.path.join(old_path, d))]\n\n#     with Pool(num_processes) as pool:\n#         for i, directory in enumerate(dirs):\n#             directory_path = os.path.join(old_path, directory)\n#             filenames_in_dir = [filename for filename in os.listdir(directory_path) if os.path.isfile(os.path.join(directory_path, filename))]\n#             sampled_filenames = random.sample(filenames_in_dir, min(number_per_class, len(filenames_in_dir)))\n\n#             new_directory_path = os.path.join(new_path, directory)\n#             os.makedirs(new_directory_path)\n\n#             args_list = [(os.path.join(directory_path, filename), os.path.join(new_directory_path, filename)) for filename in sampled_filenames]\n\n#             # Use pool.starmap for parallel execution\n#             pool.starmap(copy_files, args_list, chunksize=batch_size)\n\n#     print('Subset creation complete.')\n\n# # Example usage\n# old_val = '/kaggle/input/imagenet-object-localization-challenge/ILSVRC/Data/CLS-LOC/train'\n# new_val = '/kaggle/working/imagenet_sub_train'\n\n# make_subset(old_val, new_val, number_per_class=100)","metadata":{"_uuid":"5fe80cff-a132-4ab2-896e-29d8f9709e90","_cell_guid":"d0f51188-68fc-4756-8c6f-20452ce97907","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch_size=32\n\n#I will use a susbset of imagenetval of 10k images \nif cuda_available:\n    # path to the ImageNet dataset\n    # link --> https://www.kaggle.com/datasets/titericz/imagenet1k-val\n    \n    imagenet_data_path = '/kaggle/input/imagenet1k-subset-100k-train-and-10k-val/imagenet_subval'\n\n    # if 'imagenet_data_path' is empty, download a subset of ImageNet containing 2000 images (~250M) for test\n    if not os.path.isdir(imagenet_data_path):\n        print('%s is empty. Download a subset of ImageNet for test.' % imagenet_data_path)\n\n    print('The ImageNet dataset files are ready.')\nelse:\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')\n    \n    \n  \nif cuda_available:\n    # The following function build the data transforms for test\n    def build_val_transform(size):\n        return transforms.Compose([\n            transforms.Resize(int(math.ceil(size / 0.875))),\n            transforms.CenterCrop(size),\n            transforms.ToTensor(),\n            transforms.Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225]\n            ),\n        ])\n    \n    val_data = datasets.ImageFolder(\n            root=os.path.join(imagenet_data_path),\n            transform=build_val_transform(224)\n        )\n    \n\n    val_loader = torch.utils.data.DataLoader(\n        val_data,\n        batch_size=batch_size,  \n        shuffle = True,\n        num_workers=4,  \n        pin_memory=True,\n        drop_last=False,\n    )\n    print('The ImageNet dataloader is ready. Size : {}'.format(len(val_loader)*batch_size))\nelse:\n    data_loader = None\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')","metadata":{"_uuid":"c8f20a68-77af-4a1e-aec2-6b4f968950d3","_cell_guid":"dece162d-f096-48e6-a573-e69c2bebdba6","collapsed":false,"pycharm":{"name":"#%%\n"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now you have configured the dataset. Let's build the dataloader for evaluation.\nAgain, this will be skipped if you are in a CPU environment.","metadata":{"_uuid":"1616ad0c-4e91-419d-9339-501a56f480fb","_cell_guid":"59b9927f-0fc0-4e91-8ac4-ef6db0064ddc","trusted":true}},{"cell_type":"markdown","source":"Lets evaluate our randomly sampled network on imagenet validation set","metadata":{"_uuid":"39890b5e-7794-4839-af6c-822eca266fe1","_cell_guid":"044134e6-a76c-49e8-bf41-9622c3e711e2","trusted":true}},{"cell_type":"code","source":"from ofa.utils.common_tools import *\ndef evaluate_sub(net, data_loader=val_loader ,device=\"cuda:0\"):\n    if \"cuda\" in device:\n        net = torch.nn.DataParallel(net).to(device)\n    else:\n        net = net.to(device)\n\n    criterion = nn.CrossEntropyLoss().to(device)\n\n    net.eval()\n    net = net.to(device)\n    losses = AverageMeter()\n    top1 = AverageMeter()\n    top5 = AverageMeter()\n\n    with torch.no_grad():\n        with tqdm(total=len(data_loader), desc=\"Validate\") as t:\n            for i, (images, labels) in enumerate(data_loader):\n                images, labels = images.to(device), labels.to(device)\n                # compute output\n                output = net(images)\n                loss = criterion(output, labels)\n                # measure accuracy and record loss\n                acc1, acc5 = accuracy(output, labels, topk=(1, 5))\n\n                losses.update(loss.item(), images.size(0))\n                top1.update(acc1[0].item(), images.size(0))\n                top5.update(acc5[0].item(), images.size(0))\n                t.set_postfix(\n                    {\n                        \"loss\": losses.avg,\n                        \"top1\": top1.avg,\n                        \"top5\": top5.avg,\n                        \"img_size\": images.size(2),\n                    }\n                )\n                t.update(1)\n\n    print(\n        \"Results: loss=%.5f,\\t top1=%.1f,\\t top5=%.1f\"\n        % (losses.avg, top1.avg, top5.avg)\n    )\n    return top1.avg","metadata":{"_uuid":"6c1a6348-670d-4a1a-90f4-f23377d769fc","_cell_guid":"a67645ff-1f32-4315-91dc-fe16a373e9bc","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# %%script echo skipping\nif cuda_available:\n    top1 = evaluate_sub(subnet2)\n\n    print('Finished evaluating the pretrained sub-network: %s!' % top1)\nelse:\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')","metadata":{"_uuid":"a2af5e96-d813-4c81-be3d-47c83e0365c7","_cell_guid":"34509351-3c15-4bf8-8dbd-3f49dbf7c449","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets replace SE blocks of the network as their current implementation on OFA is not supported on FPGA so they will be replaced by pytorch SE blocks","metadata":{"_uuid":"d6c2941e-c512-4664-91cd-49ef33a0b300","_cell_guid":"5c3c8c6a-e8c5-4e7a-8bb9-19e6668ed3cb","trusted":true}},{"cell_type":"code","source":"from fpga_utils.replace_se import replace_all\n\nreplace_all(subnet2)\ntop1 = evaluate_sub(subnet2)","metadata":{"_uuid":"1dea7847-2d5c-4ca1-84d5-260727a5dbc3","_cell_guid":"76353e2d-e3d3-4b79-9cb7-ca5ffb0af938","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Accuracy Predictor \n\nThe key components of very fast neural network deployment are **accuracy predictors** and **efficiency predictors**.\nFor the accuracy predictor, it predicts the Top-1 accuracy of a given sub-network on a **holdout validation set**\n(different from the official 50K validation set) so that we do **NOT** need to run very costly inference on ImageNet\nwhile searching for specialized models. Such an accuracy predictor is trained using an accuracy dataset built with the OFA network.","metadata":{"_uuid":"e92941b1-2085-47d7-b7e9-b95357df7005","_cell_guid":"8b6a4d11-ad54-4a39-b51e-0c9659916087","trusted":true}},{"cell_type":"code","source":"# accuracy predictor\naccuracy_predictor = AccuracyPredictor(\n    pretrained=True,\n    device='cuda:0' if cuda_available else 'cpu'\n)\n\nprint('The accuracy predictor is ready!')\nprint(accuracy_predictor.model)","metadata":{"_uuid":"614ed3b7-018f-477c-8618-f87ce0924269","_cell_guid":"f68fc7af-62e5-46d1-940f-93252791d058","collapsed":false,"pycharm":{"name":"#%%\n"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets try predicting accuracy of our randomly sampled subnet","metadata":{"_uuid":"5673442d-9c37-48a6-845c-f84e6b2d6718","_cell_guid":"73a12c05-0e56-49d8-a7b8-0f0bed3e7efb","trusted":true}},{"cell_type":"code","source":"cfg['r']= [224]\n\nacc = accuracy_predictor.predict_accuracy([cfg])\nprint(acc*100)","metadata":{"_uuid":"8960373f-1889-4628-93e4-227807439787","_cell_guid":"610d4f79-a269-4cf0-bd4b-d09e9e26e7f0","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, we have the powerful **accuracy predictor**. We then introduce two types of **efficiency predictors**: the latency predictor and the FLOPs predictor. \n\nThe intuition of having efficiency predictors, especially the latency predictor, is that measuring the latency of a sub-network on-the-fly is also costly, especially for FPGA devices becuase it takes hours of work to implement even a single network and then measure latency.\n\nThe latency predictor is designed to eliminate this cost.","metadata":{"_uuid":"fe970f2f-26c4-4366-80cd-62fc9d5a7814","_cell_guid":"8f7f519e-9474-4816-96f6-352e360eec2e","trusted":true}},{"cell_type":"markdown","source":"## 5. Searching with Latency Constraint\nNow, let's proceed towards searching for efficient network under latency constraint. We use the same accuracy predictor since accuracy predictors are agnostic to the types of efficiency constraint . For the efficiency predictor, we build a latency table.","metadata":{"_uuid":"6b4e342f-e521-4ef6-aaf1-52500decf8db","_cell_guid":"84c3dd41-3330-4748-9a91-efe3fc31efef","trusted":true}},{"cell_type":"code","source":"class LatencyTable():\n    def __init__ (self, path = 'datasets/latency_lut.npy'):\n        self.path = path \n        self.efficiency_dict = np.load(self.path, allow_pickle=True).item()\n        \n        \n    #exception -- I have not included the latency of average pooling \n    def predict_efficiency(self, sample):\n        input_size = sample.get(\"r\", [224])\n        input_size = input_size[0]\n        assert \"ks\" in sample and \"e\" in sample and \"d\" in sample\n        assert len(sample[\"ks\"]) == len(sample[\"e\"]) and len(sample[\"ks\"]) == 20\n        assert len(sample[\"d\"]) == 5\n        \n        total_latency = 0 \n        \n        for i in range(20):\n            stage = i // 4\n            depth_max = sample[\"d\"][stage]\n            depth = i % 4 + 1\n            if depth > depth_max:\n                continue\n            ks, e = sample[\"ks\"][i], sample[\"e\"][i]\n            \n            \n            total_latency+= self.efficiency_dict[\"mobile_inverted_blocks\"][i+1][(ks, e)]\n            \n\n\n        for key in self.efficiency_dict[\"other_blocks\"]:\n            total_latency+=self.efficiency_dict[\"other_blocks\"][key]\n            \n        \n        return total_latency\n    \n        \n\nlatency_estimator = LatencyTable()\n\nprint('The Latency lookup table is ready!')","metadata":{"_uuid":"ce28329d-35e7-4796-9541-7dd09c86b1cb","_cell_guid":"62bd0418-2cd8-4b1d-942d-caf9ec792c4b","collapsed":false,"pycharm":{"is_executing":true,"name":"#%%\n"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets try to find latency of our random sampled subnet","metadata":{"_uuid":"3fc31f74-7d40-407e-b08c-cbb951788b94","_cell_guid":"75a3d901-3733-4d36-8f48-3c69c9581fea","trusted":true}},{"cell_type":"code","source":"lat = latency_estimator.predict_efficiency(cfg)   #as it returns bytes/ops\nprint('Estimated Latency = {} ms'.format(lat))","metadata":{"_uuid":"fbcccb12-9828-477e-8a37-b7d8bdac7d55","_cell_guid":"7c83a430-afd2-4595-963b-f335d8b5439b","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets try lowest Depth network and highest depth network latency","metadata":{"_uuid":"3dec93ef-bc7d-4f9e-90f7-0ebe949054be","_cell_guid":"5e74820c-ff49-4b8c-87cb-6e94069a13c1","trusted":true}},{"cell_type":"code","source":"max_cfg = ofa_network.set_max_net()\nmax_net = ofa_network.get_active_subnet()\n\nmin_cfg = ofa_network.set_active_subnet(ks=3, e=4, d=2)\nmin_net = ofa_network.get_active_subnet()","metadata":{"_uuid":"e74344c6-7a87-43c8-9d92-5d2558051d4a","_cell_guid":"41e4c729-3ad6-41e4-b22e-dac1e69684bb","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lat1 = latency_estimator.predict_efficiency(max_cfg)   #as it returns bytes/ops\nprint('Estimated Latency Max= {} ms'.format(lat1))\n\nlat2 = latency_estimator.predict_efficiency(min_cfg)   #as it returns bytes/ops\nprint('Estimated Latency Min = {} ms'.format(lat2))","metadata":{"_uuid":"8aecbbba-3e67-46e0-9fef-475afaf5f73e","_cell_guid":"55854701-5e1e-4283-bf5a-01d8ec8be5ff","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nplt.title('Variation in Latency in the search Space')\nplt.bar(['Shallow Network', 'Widest Network'], [lat2,lat1], color = ['green', 'red'], width = 0.4)\nplt.ylabel('Latency(ms)')\nplt.ylim(0, 60)\nplt.show()","metadata":{"_uuid":"38b3755b-264f-4bcc-87da-79d415814da7","_cell_guid":"63a60981-e35b-4bde-bf9d-d2bd72a5f740","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Run Evolutionary search","metadata":{"_uuid":"7f17cfda-4972-4f1b-8e7a-40b469f0e541","_cell_guid":"2c3dde1c-20cd-41af-b9ce-aeace491a287","trusted":true}},{"cell_type":"markdown","source":"Lets try these constraint once","metadata":{"_uuid":"1df4bde8-21a2-4799-892c-80e8d8512e3c","_cell_guid":"7e2d5ea2-772d-42f0-863a-61874782a32f","trusted":true}},{"cell_type":"code","source":"#  Hyper-parameters for the evolutionary search process\n\nP = 10000  # The size of population in each generation\nN = 1200  # How many generations of population to be searched\nr = 0.3  # The ratio of networks that are used as parents for next generation\n\nparams = {\n    'constraint_type': 'latency', # Let's do latency constrained search\n    'efficiency_constraint': 13,  # latency constraint , suggested range [10, 45]\n    'mutate_prob': 0.3, # The probability of mutation in evolutionary search\n    'mutation_ratio': 0.4, # The ratio of networks that are generated through mutation in generation n >= 2.\n    'efficiency_predictor': latency_estimator, # To use a predefined efficiency predictor.\n    'accuracy_predictor': accuracy_predictor, # To use a predefined accuracy_predictor predictor.\n    'population_size': P,\n    'max_time_budget': N,\n    'parent_ratio': r,\n}\n\n\nfinder = EvolutionFinder(**params)\n\nlatency_list = [40, 32, 25, 18, 12, 11, 10]\n\nresult_lis = []\nresult_valids = []\ninfo = []\nfor latency in latency_list:\n    st = time.time()\n    finder.set_efficiency_constraint(latency)\n    best_valids, best_info = finder.run_evolution_search()\n    ed = time.time()\n    \n    print('Found best architecture at latency <= %.2f ms in %.2f seconds! It achieves %.2f%s predicted accuracy with latency of %.2f ms./n' % (latency, ed-st, best_info[0] * 100, '%',best_info[-1]))\n    result_lis.append(best_info)\n    result_valids.append(best_valids)\n    info.append(ed-st)","metadata":{"_uuid":"337a45a5-262b-4d71-a67b-684f3d4deb79","_cell_guid":"fa0e53fb-4bd0-4030-9e7b-9043a8206a41","collapsed":false,"pycharm":{"name":"#%%\n"},"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The results after runnint them are !\n\nLets visualize and evalute the found architecutre","metadata":{"_uuid":"915a9451-60c8-4a1d-a5d4-a6157e47b2a6","_cell_guid":"a123c81b-3a27-4aa5-9e57-98960e64985f","trusted":true}},{"cell_type":"code","source":"!mkdir Models","metadata":{"_uuid":"1374eab8-6211-445d-b5ab-f44af98757b7","_cell_guid":"d76ade3e-8656-4e04-9a49-fca08df3b51c","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7. Visualizing and Evaluating Searched Models\nLets build model from config and Replace SE blocks in all these networks","metadata":{"_uuid":"a89eb086-999e-4f0b-83d3-e206a80f0da8","_cell_guid":"5c8cfbf1-5cb8-498f-8110-b6efa3fa7f92","trusted":true}},{"cell_type":"code","source":"print('Replacing SE blocks of OFA and \\n Saving the searched models ')\nfor arch in result_lis:\n    cfg = arch[1]\n    visualize_subnet(cfg)  \n    ofa_network.set_active_subnet(cfg['ks'], cfg['e'], cfg['d'])\n    network = ofa_network.get_active_subnet()\n    acc = evaluate_sub(network)\n    \n    print('The evaluated accuracy is : {}'.format(acc))\n    print('--'*50)\n    \nmodels = {}   # latency, network pair\nfor i, arch in enumerate(result_lis):\n    cfg = arch[1]\n    ofa_network.set_active_subnet(cfg['ks'], cfg['e'], cfg['d'])\n    network = ofa_network.get_active_subnet(preserve_weight=True)\n    models[new_latency_list[i]] = network\n\nfor key, model in models.items():\n    replace_all(model)\n    name = f\"Models/model_search_{key}.pth\"\n    torch.save(model, name)\n    print('Done')","metadata":{"_uuid":"3926e24e-3ff7-469f-8476-4f529d433cfa","_cell_guid":"b320c23f-93d8-4ee3-93bd-a1765f279b2a","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Saving architecture configuration to a File**","metadata":{"_uuid":"6b55b809-8ea2-4f07-9f6a-7fec3def684b","_cell_guid":"6900c50d-505a-4725-9198-7a556049821b","trusted":true}},{"cell_type":"code","source":"import json\n\ndata = {\n    \"arch\": result_lis,\n    \"info\": info\n}\n\nwith open('search_data.json', 'w') as file:\n    json.dump(data, file)\n    \nprint('The cofig data is exported, \\n The exported Data is:')\nprint(data)","metadata":{"_uuid":"b3c4e670-fd9a-4101-9026-857c8eaa6ddc","_cell_guid":"3c5e3db6-6649-4620-b395-103904eec52b","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 8. Retrain the Searched network\nLets fine tune the network on a subset of training set of imagenet.","metadata":{"_uuid":"8e200e47-208e-4c1a-9284-c02e0ed70a0b","_cell_guid":"afbbdb50-bf35-4170-97fd-fa2f2355b0c0","trusted":true}},{"cell_type":"markdown","source":"I previously performed FPGA aware Neural architecutral search and found different models with different latency through evolutionary algorithm.\n\nThe obtained network from search process will be now retrained to improve accuracy. Also these models wt were intialized with the help of OFA!","metadata":{"_uuid":"64743ae8-1816-49bc-8d80-b261d4423cfb","_cell_guid":"82bd3e62-0f92-4d26-9991-ce9c8d26ce2b","trusted":true}},{"cell_type":"code","source":"import torch \nimport torchvision\nimport os\nimport torch.nn as nn\nfrom torchvision import transforms, datasets\nimport math\nimport time\nfrom tqdm import tqdm\nimport shutil\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"33432336-296b-464f-bd14-6e39ea995376","_cell_guid":"a2e3a91a-8119-496c-85c6-df09ac11a018","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cuda_available = torch.cuda.is_available()\nif cuda_available:\n    torch.backends.cudnn.enabled = True\n    torch.backends.cudnn.benchmark = True\n    print('Using GPU.')\nelse:\n    print('Using CPU.')","metadata":{"_uuid":"9f60a22d-8932-4b20-909f-0f72e11d61a2","_cell_guid":"995ad112-0da3-453d-981c-7ffb3dac67d0","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"batch_size=32\n\n#I will use a susbset of imagenetval of 10k images \nif cuda_available:\n    # path to the ImageNet dataset\n    # link --> https://www.kaggle.com/datasets/titericz/imagenet1k-val\n    \n    imagenet_data_path = '/kaggle/input/imagenet1k-subset-100k-train-and-10k-val/imagenet_subval'\n\n    # if 'imagenet_data_path' is empty, download a subset of ImageNet containing 2000 images (~250M) for test\n    if not os.path.isdir(imagenet_data_path):\n        print('%s is empty. Download a subset of ImageNet for test.' % imagenet_data_path)\n\n    print('The ImageNet dataset files are ready.')\nelse:\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')\n    \n    \n  \nif cuda_available:\n    # The following function build the data transforms for test\n    def build_val_transform(size):\n        return transforms.Compose([\n            transforms.Resize(int(math.ceil(size / 0.875))),\n            transforms.CenterCrop(size),\n            transforms.ToTensor(),\n            transforms.Normalize(\n                mean=[0.485, 0.456, 0.406],\n                std=[0.229, 0.224, 0.225]\n            ),\n        ])\n    \n    val_data = datasets.ImageFolder(\n            root=os.path.join(imagenet_data_path),\n            transform=build_val_transform(224)\n        )\n    \n\n    val_loader = torch.utils.data.DataLoader(\n        val_data,\n        batch_size=batch_size,  \n        shuffle = True,\n        num_workers=4,  \n        pin_memory=True,\n        drop_last=False,\n    )\n    print('The ImageNet dataloader is ready. Size : {}'.format(len(val_loader)*batch_size))\nelse:\n    data_loader = None\n    print('Since GPU is not found in the environment, we skip all scripts related to ImageNet evaluation.')","metadata":{"_uuid":"a6cdacee-d402-4b3e-9ac4-fa080dd481c0","_cell_guid":"b49ecee7-e265-4267-9cb1-020ef5a8d2cb","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ntrain_path = '/kaggle/input/imagenet1k-subset-100k-train-and-10k-val/imagenet_subtrain'\n\ntrain_data = datasets.ImageFolder(\n            root= train_path,\n            transform=build_val_transform(224)\n        )\n\ntrain_loader = torch.utils.data.DataLoader(\n        train_data,\n        batch_size=batch_size, \n        shuffle = True,\n        num_workers=4,  \n        pin_memory=True,\n        drop_last=False,\n    )\n\nprint('The ImageNet train set is ready. Size : {}'.format(len(train_loader)*batch_size))","metadata":{"_uuid":"36496cc1-9614-4f63-81d1-ee99e583cc2e","_cell_guid":"a6f8888b-b40a-4741-9781-62dd5ca523d0","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataloaders = {}\ndataloaders['train'] = train_loader\ndataloaders['val'] = val_loader\n\ndataset_sizes = {'train': len(train_loader)*32,\n                'val': len(val_loader)*32}\nprint(dataset_sizes)","metadata":{"_uuid":"e3d57b25-eae9-4430-a9c7-9f49c8dd55ad","_cell_guid":"b768509f-197f-4991-8013-413fbaedaf22","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def train_model(model, criterion, optimizer, scheduler, num_epochs=5):\n    since = time.time()\n\n    #storing epoch data\n    epoch_data = {\n        'epoch': [],\n        'train': {'loss': [], 'top1_acc': [], 'top5_acc': []},\n        'val': {'loss': [], 'top1_acc': [], 'top5_acc': []}\n    }\n    \n    # Create a temporary directory\n    tempdir = '/kaggle/working/temp'\n    os.makedirs(tempdir, exist_ok=True)\n    best_model_params_path = os.path.join(tempdir, 'best_model_params.pt')\n\n    torch.save(model.state_dict(), best_model_params_path)\n    best_acc = 0.0\n\n    for epoch in range(num_epochs):\n        print(f'Epoch {epoch+1}/{num_epochs}')\n        print('-' * 10)\n        epoch_data['epoch'].append(epoch+1)\n        \n        for phase in ['train', 'val']:\n            if phase == 'train':\n                model.train()\n            else:\n                model.eval()\n            running_loss = 0.0\n            top1_corrects = 0\n            top5_corrects = 0\n            \n            for inputs, labels in tqdm(dataloaders[phase], leave=False):\n                inputs = inputs.to(device)\n                labels = labels.to(device)\n\n                optimizer.zero_grad()\n\n                with torch.set_grad_enabled(phase == 'train'):\n                    outputs = model(inputs)\n                    _, preds = torch.max(outputs, 1)\n                    loss = criterion(outputs, labels)\n\n                    if phase == 'train':\n                        loss.backward()\n                        optimizer.step()\n\n                running_loss += loss.item() * inputs.size(0)\n                \n                # Calculate top-1 accuracy\n                top1_corrects += torch.sum(preds == labels.data)\n                \n                # Calculate top-5 accuracy\n                _, top5_preds = torch.topk(outputs, 5, dim=1)\n                top5_corrects += torch.sum(top5_preds == labels.view(-1, 1))\n\n            if phase == 'train':\n                scheduler.step()\n\n            epoch_loss = running_loss / dataset_sizes[phase]\n            epoch_top1_acc = top1_corrects.double() / dataset_sizes[phase]\n            epoch_top5_acc = top5_corrects.double() / dataset_sizes[phase]\n            \n            epoch_data[phase]['loss'].append(epoch_loss)\n            epoch_data[phase]['top1_acc'].append(epoch_top1_acc)\n            epoch_data[phase]['top5_acc'].append(epoch_top5_acc)\n\n            print(f'{phase} Loss: {epoch_loss:.4f} Top-1 Acc: {epoch_top1_acc:.4f} Top-5 Acc: {epoch_top5_acc:.4f}')\n\n            if phase == 'val' and epoch_top1_acc > best_acc:\n                best_acc = epoch_top1_acc\n                best_top5 = epoch_top5_acc\n                torch.save(model.state_dict(), best_model_params_path)\n\n        print()\n\n    time_elapsed = time.time() - since\n    print(f'Training complete in {time_elapsed // 60:.0f}m {time_elapsed % 60:.0f}s')\n    print(f'Best val Top-1 Acc {best_acc:4f} /n Best val Top-5: {best_top5:4f}')\n\n    model.load_state_dict(torch.load(best_model_params_path))\n\n    # Clean up the temporary directory\n    shutil.rmtree(tempdir)\n\n    return model, epoch_data\n","metadata":{"_uuid":"ec1f38ea-c450-466b-b68f-8b69452a9434","_cell_guid":"9c1f3261-c421-412c-832e-0430781f8477","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.cuda.empty_cache()\nmodel = torch.load('Models/model_search_11.pth')\nmodel = model.to(device)\n\ncriterion = nn.CrossEntropyLoss()\n# Observe that all parameters are being optimized\noptimizer_ft = torch.optim.SGD(model.parameters(), lr=0.00001, momentum=0.9)\n\n# Decay LR by a factor of 0.1 every 7 epochs\nexp_lr_scheduler = torch.optim.lr_scheduler.StepLR(optimizer_ft, step_size=2, gamma=0.5)\n\nmodel, epoch_data = train_model(model, criterion, optimizer_ft, exp_lr_scheduler,\n                   num_epochs=2)\n\ntorch.save(model, 'retrained_model_search_11.pth')\nprint('*****************************************************')\nprint(epoch_data)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### It is clear that accuracy is recoverable and we can obtain comparable accuracy with our efficiency contraints by retraining for few epochs","metadata":{"_uuid":"1e7aec39-a70e-4250-b393-e5c5e2476e8b","_cell_guid":"ff3894d4-cec7-4720-b42f-d6d594dcc950","trusted":true}},{"cell_type":"markdown","source":"## 6. Insights and Comparision","metadata":{"_uuid":"739e1062-7782-450e-961b-e33441bd4165","_cell_guid":"50055b45-a44f-4f89-8712-69e0f9c96819","trusted":true}},{"cell_type":"code","source":"# 'ResNet-152':81.3,\n#  'Efficientnet V2 Small': 83.6,","metadata":{"_uuid":"62cdea74-7374-4e1c-ad22-2f36c4ee1197","_cell_guid":"1c1735e3-d8f1-4274-9e6c-6ffc2744c0b7","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import shutil\n\n# def zip_folder(folder_path, zip_path):\n#     try:\n#         # Zip the entire folder and its contents\n#         shutil.make_archive(zip_path, 'zip', folder_path)\n#         print(f'Folder \"{folder_path}\" successfully zipped to \"{zip_path}.zip\"')\n#     except Exception as e:\n#         print(f'Error zipping folder: {e}')\n\n# # Example usage:\n\n# zip_folder('/kaggle/working/Models/', '/kaggle/working/models')","metadata":{"_uuid":"4d002645-9da2-4470-8896-ddbc092a283a","_cell_guid":"7e3d79eb-eb5c-4365-b914-9b60a52247f6","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# archs = [\n# ['Efficientnet B0', 77.8, 19.5],\n# ['Mobilenet V3_large', 77, 11.5],\n# ['nvidia_efficientnet_b4',82, 20.3]\n# ]\n\n\n# y = [res[0]*100 for res in result_lis]+ [79.5] #added evaulated accuracy after one epcoh train\n# x = [1/res[2] for res in result_lis]+ [20.1]\n# plt.figure(figsize=(12, 8))\n# plt.xlabel('Arithmetic Intensity (ops/byte)', fontsize=14)\n# plt.ylabel('Predicted Accuracy', fontsize=14)\n\n# # Plot the line and scatter plot with 'X' marker for each point in the same color\n# plt.plot(x, y, marker='X', linestyle='-', color='blue', markersize=8, label = 'Ours')\n# plt.plot(archs[0][2], archs[0][1], marker='*', color='red', label ='other', markersize=10 )\n# plt.plot(archs[1][2],archs[1][1], marker = '*', color = 'red', markersize = 10)\n# plt.plot(archs[2][2],archs[2][1], marker = '*', color = 'red', markersize = 10)\n\n\n# # Annotate the middle point\n# plt.text(20.2,79, f' Evaluated Accuracy \\n after 1 epoch', fontsize = 10)\n# plt.text(17, 85.2, f'      FPGA specialized networks', fontsize=12, verticalalignment='bottom', horizontalalignment='left', color='blue')\n# plt.text(archs[0][2], archs[0][1], f'  Efficientnet_b0', fontsize=12 )\n# plt.text(archs[1][2],archs[1][1], f' {archs[1][0]}')\n# plt.text(archs[2][2]+0.1,archs[2][1], f' {archs[2][0]}', fontsize=12)\n\n# plt.ylim(75, 89)\n# plt.xlim(10, 24)\n# # Add legend\n# plt.legend(fontsize=12)\n# plt.title('Searching for FPGAs')\n# # Show the plot\n# plt.show()","metadata":{"_uuid":"3264d105-89ed-4d3a-9ee1-51f3e4a0c946","_cell_guid":"6a2e9b32-330b-4478-b0d9-82f6c74b763e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"_uuid":"fe3c0f71-9689-49e4-a435-2bae691a423c","_cell_guid":"8d6227c6-7a8f-4332-a53b-109b39a118f0","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]}]}