{
  "id": 105254,
  "title": "Amazon SageMaker Example",
  "url": "/competitions/kuzushiji-recognition/discussion/105254",
  "author_name": "",
  "post_date": "2019-08-22T02:33:44.640390100Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For a long time I have wanted to experiment with <a href=\"https://aws.amazon.com/sagemaker/\">Amazon Sagemaker</a>, so I decided to try it on the notebook with the best score in the Notebooks section of this contest, which is this Centernet Keypoint Detector one from K_mat:  <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">https://www.kaggle.com/kmat2019/centernet-keypoint-detector</a></p>\n\n<p>SageMaker is tool from Amazon that makes using ML easier in various ways, such as being able to run a notebook from your local CPU machine but still do the training part on remote AWS GPU instances. It also has all sorts of built-in ML algorithms that are optimized to be faster than usual, including an optimized version of Tensorflow (which now also includes Keras).</p>\n\n<p>I do not have a GPU myself, so every time I need to use one I first create the notebook as usual on my AWS CPU server, do all sorts of testing, and when I am ready to train the model for real I save an AMI of server (an AMI is like a backup copy) which I then launch a copy of using an AWS GPU server. </p>\n\n<p>But, it would be nice not to have to do that, plus I have to keep checking to see when the GPU server is done with training, or else I would be paying for it while it is no longer in use. So SageMaker seemed like a good possible solution.</p>\n\n<p>I first tried the Centernet Keypoint Detector notebook on my existing AWS CPU server (not using SageMaker) and it got a submission score of 0.515, which is a little less than the original notebook, but because I was using a CPU and not a GPU like he was, it probably trained it less for the same number of epochs. Mainly though, it took around 2 days for training, which is annoyingly long.</p>\n\n<p>I installed SageMaker on my server, and after some time dealing with various configuration issues getting Boto and AWS CLI to work with my AWS account info and S3 buckets, I got SageMaker sort of working using the demo notebooks that came with it.  I am not really a programmer though, and there was one error I could not figure out, so I hired somebody to fix it for me who had experience with SageMaker. He fixed it in a matter of minutes, so I then decided to pay him to get the Centernet Keypoint Detector notebook working for me in SageMaker, because it would have taken my days to do it myself.</p>\n\n<p>It ended up taking him 7 hours, so there is no way I could have done it. He wrote: \"There was some errors for resolving out the dependencies in their training docker, thats where most time was consumed, but it was one time work, Now we have solution for that and can handle that from now on in much less time. It's the total time for building the understanding what is being done in the notebook to running the training job.\"   And also:   \"...issue was related to external libraries like cv2, glob etc. There are workarounds for that to install those, One is implemented and you can use those as placeholders for next projects\".</p>\n\n<p>What I mainly wanted to find out is if using SageMaker (with a GPU) for it was any faster than running the original notebook K_mat posted at:  <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">https://www.kaggle.com/kmat2019/centernet-keypoint-detector</a>.  I am not sure what GPU K-mat used, but with SageMaker I used an AWS V100 which is super fast, so I doubt his was any faster.  For comparison purposes, in step 1 of the training part of his notebook, SageMaker took 801s (9s/step) for the 1st epoch.  K_mat's notebook was much faster at 439s (5s/step).  Most interestingly, my 8 core AWS CPU server (m4.xlarge) took 1844s (20s/step) which was not that much worse, considering how much cheaper it is to use a CPU.</p>\n\n<p>I did not let the SageMaker training finish, so I don't know what the final accuracy would have been, but I have no reason to think it would have been any better than the original notebook. All of this was just about trying to make things either easier or faster for me, but in the end it was not worth it for what I am doing.</p>\n\n<p>Something to keep in mind though is that many people could use SageMaker via the AWS management console (similar to how Kaggle offers Kaggle kernals for you to use), instead of using it locally in script mode. This way, you create the notebook directly in the AWS cloud without running any server yourself.  That might have made it so none of those errors happened, but it still would have involved all the same programming to convert the Kaggle notebook to SageMaker format.</p>\n\n<p>In case anybody is interested, below is all the code I used. SageMaker script works using 2 files. First, it uses the original notebook code as a .py file, with a few simple lines of code added for SageMaker. This .py file is then run as part of a SageMaker notebook, which tells it where your data is stored and which type of AWS instance to use for the training.  </p>\n\n<p>Here's the code:</p>\n\n<p>================================</p>\n\n<p>sagemaker_centernet_script.py</p>\n\n<p>import argparse, os\nos.system(\"apt-get update\")\nos.system(\"apt-get install -y libsm6 libxext6 libxrender-dev\")\nos.system(\"pip install opencv-python-headless\")</p>\n\n<p>os.system(\"pip install Pillow\")\nos.system(\"pip install matplotlib\")\nos.system(\"pip install glob3\")</p>\n\n<p>import numpy as np\nimport json\nimport pandas as pd\nfrom PIL import Image, ImageDraw\nimport matplotlib.pyplot as plt\nfrom pandas.io.json import json_normalize\nimport random\nimport tensorflow as tf\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import KFold,train_test_split\nimport matplotlib.pyplot as plt\nimport glob\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.layers import Dense,Dropout, Conv2D,Conv2DTranspose, BatchNormalization, Activation,AveragePooling2D,GlobalAveragePooling2D, Input, Concatenate, MaxPool2D, Add, UpSampling2D, LeakyReLU,ZeroPadding2D\nfrom keras.models import Model\nfrom keras.objectives import mean_squared_error\nfrom keras import backend as K\nfrom keras.losses import binary_crossentropy\nfrom keras.callbacks import ModelCheckpoint, EarlyStopping, TensorBoard, ReduceLROnPlateau,LearningRateScheduler</p>\n\n<p>from keras.optimizers import Adam, RMSprop, SGD</p>\n\n<p>import boto3 </p>\n\n<p>import keras.backend.tensorflow_backend as K\nK.set_session</p>\n\n<p>category_n=1\nimport cv2\ninput_width,input_height=512, 512</p>\n\n<p>def Datagen_sizecheck_model(filenames, batch_size, size_detection_mode=True, is_train=True,random_crop=True):\n  x=[]\n  y=[]</p>\n\n<p>count=0</p>\n\n<p>while True:\n    for i in range(len(filenames)):\n      if random_crop:\n        crop_ratio=np.random.uniform(0.7,1)\n      else:\n        crop_ratio=1\n      with Image.open(filenames[i][0]) as f:\n        #random crop\n        if random_crop and is_train:\n          pic_width,pic_height=f.size\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8)\n          top_offset=np.random.randint(0,pic_height-int(crop_ratio*pic_height))\n          left_offset=np.random.randint(0,pic_width-int(crop_ratio*pic_width))\n          bottom_offset=top_offset+int(crop_ratio*pic_height)\n          right_offset=left_offset+int(crop_ratio*pic_width)\n          f=cv2.resize(f[top_offset:bottom_offset,left_offset:right_offset,:],(input_height,input_width))\n        else:\n          f=f.resize((input_width, input_height))\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8) <br>\n        x.append(f)</p>\n\n<pre><code>  if random_crop and is_train:\n    y.append(filenames[i][1]-np.log(crop_ratio))\n  else:\n    y.append(filenames[i][1])\n\n  count+=1\n  if count==batch_size:\n    x=np.array(x, dtype=np.float32)\n    y=np.array(y, dtype=np.float32)\n\n    inputs=x/255\n    targets=y       \n    x=[]\n    y=[]\n    count=0\n    yield inputs, targets\n</code></pre>\n\n<p>def aggregation_block(x_shallow, x_deep, deep_ch, out_ch):\n  x_deep= Conv2DTranspose(deep_ch, kernel_size=2, strides=2, padding='same', use_bias=False)(x_deep)\n  x_deep = BatchNormalization()(x_deep) <br>\n  x_deep = LeakyReLU(alpha=0.1)(x_deep)\n  x = Concatenate()([x_shallow, x_deep])\n  x=Conv2D(out_ch, kernel_size=1, strides=1, padding=\"same\")(x)\n  x = BatchNormalization()(x) <br>\n  x = LeakyReLU(alpha=0.1)(x)\n  return x</p>\n\n<p>def cbr(x, out_layer, kernel, stride):\n  x=Conv2D(out_layer, kernel_size=kernel, strides=stride, padding=\"same\")(x)\n  x = BatchNormalization()(x)\n  x = LeakyReLU(alpha=0.1)(x)\n  return x</p>\n\n<p>def resblock(x_in,layer_n):\n  x=cbr(x_in,layer_n,3,1)\n  x=cbr(x,layer_n,3,1)\n  x=Add()([x,x_in])\n  return x  </p>\n\n<h1>I use the same network at CenterNet</h1>\n\n<p>def create_model(input_shape, size_detection_mode=True, aggregation=True):\n    input_layer = Input(input_shape)</p>\n\n<pre><code>#resized input\ninput_layer_1=AveragePooling2D(2)(input_layer)\ninput_layer_2=AveragePooling2D(2)(input_layer_1)\n\n#### ENCODER ####\n\nx_0= cbr(input_layer, 16, 3, 2)#512-&gt;256\nconcat_1 = Concatenate()([x_0, input_layer_1])\n\nx_1= cbr(concat_1, 32, 3, 2)#256-&gt;128\nconcat_2 = Concatenate()([x_1, input_layer_2])\n\nx_2= cbr(concat_2, 64, 3, 2)#128-&gt;64\n\nx=cbr(x_2,64,3,1)\nx=resblock(x,64)\nx=resblock(x,64)\n\nx_3= cbr(x, 128, 3, 2)#64-&gt;32\nx= cbr(x_3, 128, 3, 1)\nx=resblock(x,128)\nx=resblock(x,128)\nx=resblock(x,128)\n\nx_4= cbr(x, 256, 3, 2)#32-&gt;16\nx= cbr(x_4, 256, 3, 1)\nx=resblock(x,256)\nx=resblock(x,256)\nx=resblock(x,256)\nx=resblock(x,256)\nx=resblock(x,256)\n\nx_5= cbr(x, 512, 3, 2)#16-&gt;8\nx= cbr(x_5, 512, 3, 1)\n\nx=resblock(x,512)\nx=resblock(x,512)\nx=resblock(x,512)\n\nif size_detection_mode:\n  x=GlobalAveragePooling2D()(x)\n  x=Dropout(0.2)(x)\n  out=Dense(1,activation=\"linear\")(x)\n\nelse:#centernet mode\n#### DECODER ####\n  x_1= cbr(x_1, output_layer_n, 1, 1)\n  x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n  x_2= cbr(x_2, output_layer_n, 1, 1)\n  x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n  x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n  x_3= cbr(x_3, output_layer_n, 1, 1)\n  x_3 = aggregation_block(x_3, x_4, output_layer_n, output_layer_n) \n  x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n  x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n\n  x_4= cbr(x_4, output_layer_n, 1, 1)\n\n  x=cbr(x, output_layer_n, 1, 1)\n  x= UpSampling2D(size=(2, 2))(x)#8-&gt;16 tconvのがいいか\n\n  x = Concatenate()([x, x_4])\n  x=cbr(x, output_layer_n, 3, 1)\n  x= UpSampling2D(size=(2, 2))(x)#16-&gt;32\n\n  x = Concatenate()([x, x_3])\n  x=cbr(x, output_layer_n, 3, 1)\n  x= UpSampling2D(size=(2, 2))(x)#32-&gt;64   128のがいいかも？ \n\n  x = Concatenate()([x, x_2])\n  x=cbr(x, output_layer_n, 3, 1)\n  x= UpSampling2D(size=(2, 2))(x)#64-&gt;128 \n\n  x = Concatenate()([x, x_1])\n  x=Conv2D(output_layer_n, kernel_size=3, strides=1, padding=\"same\")(x)\n  out = Activation(\"sigmoid\")(x)\n\nmodel=Model(input_layer, out)\n\nreturn model\n</code></pre>\n\n<p>def model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size=32):\n    hist = model.fit_generator(\n        Datagen_sizecheck_model(train_list,batch_size, is_train=True,random_crop=True),\n        steps_per_epoch = len(train_list) // batch_size,\n        epochs = n_epoch,\n        validation_data=Datagen_sizecheck_model(cv_list,batch_size, is_train=False,random_crop=False),\n        validation_steps = len(cv_list) // batch_size,\n        callbacks = [lr_schedule, model_checkpoint],#[early_stopping, reduce_lr, model_checkpoint],\n        shuffle = True,\n        verbose = 1\n    )\n    return hist</p>\n\n<p>def lrs(epoch):\n    lr = 0.0005\n    if epoch&gt;10:\n        lr = 0.0001\n    return lr</p>\n\n<p>if <strong>name</strong> == '<strong>main</strong>':</p>\n\n<pre><code>parser = argparse.ArgumentParser()\n\nparser.add_argument('--epochs', type=int, default=10)\nparser.add_argument('--learning-rate', type=float, default=0.01)\nparser.add_argument('--batch-size', type=int, default=128)\nparser.add_argument('--gpu-count', type=int, default=os.environ['SM_NUM_GPUS'])\nparser.add_argument('--model-dir', type=str, default=os.environ['SM_MODEL_DIR'])\nparser.add_argument('--training', type=str, default=os.environ['SM_CHANNEL_TRAINING'])\nparser.add_argument('--validation', type=str, default=os.environ['SM_CHANNEL_VALIDATION'])\n\nargs, _ = parser.parse_known_args()\ns3 = boto3.client('s3')\n\nepochs     = args.epochs\nlr         = args.learning_rate\nbatch_size = args.batch_size\ngpu_count  = args.gpu_count\nmodel_dir  = args.model_dir\ntraining_dir   = args.training\nvalidation_dir = args.validation\n\npath_1= \"/opt/ml/input/data/training/train_csv/train.csv\"\npath_2=\"/opt/ml/input/data/training/\"\npath_3=\"/opt/ml/input/data/validation/\"\npath_4=\"/opt/ml/input/data/validation/submission/sample_submission.csv\"\nfinal_weights_step1 = \"/opt/ml/input/data/training/final_weights/final_weights_step1.hdf5\"\n\n\ndf_train=pd.read_csv(path_1)\n#print(df_train.head())\n#print(df_train.shape)\ndf_train=df_train.dropna(axis=0, how='any')#you can use nan data(page with no letter)\ndf_train=df_train.reset_index(drop=True)\n#print(df_train.shape)\n\nannotation_list_train=[]\ncategory_names=set()\n\nfor i in range(len(df_train)):\n    ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,x,y,width,height for each picture\n    category_names=category_names.union({i for i in ann[:,0]})\n\ncategory_names=sorted(category_names)\ndict_cat={list(category_names)[j]:str(j) for j in range(len(category_names))}\ninv_dict_cat={str(j):list(category_names)[j] for j in range(len(category_names))}\n#print(dict_cat)\n\nfor i in range(len(df_train)):\n    ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,left,top,width,height for each picture\n    for j,category_name in enumerate(ann[:,0]):\n        ann[j,0]=int(dict_cat[category_name])  \n    ann=ann.astype('int32')\n    ann[:,1]+=ann[:,3]//2#center_x\n    ann[:,2]+=ann[:,4]//2#center_y\n    annotation_list_train.append([\"{}{}.jpg\".format(path_2,df_train.loc[i,\"image_id\"]),ann])\n\ndf_submission=pd.read_csv(path_4)\nid_test=path_3+df_submission[\"image_id\"]+\".jpg\"\n\naspect_ratio_pic_all=[]\naspect_ratio_pic_all_test=[]\naverage_letter_size_all=[]\ntrain_input_for_size_estimate=[]\nresize_dir=\"resized/\"\nif os.path.exists(resize_dir) == False:os.mkdir(resize_dir)\nfor i in range(len(annotation_list_train)):\n    with Image.open(annotation_list_train[i][0]) as f:\n        width,height=f.size\n        area=width*height\n        aspect_ratio_pic=height/width\n        aspect_ratio_pic_all.append(aspect_ratio_pic)\n        letter_size=annotation_list_train[i][1][:,3]*annotation_list_train[i][1][:,4]\n        letter_size_ratio=letter_size/area\n\n        average_letter_size=np.mean(letter_size_ratio)\n        average_letter_size_all.append(average_letter_size)\n        train_input_for_size_estimate.append([annotation_list_train[i][0],np.log(average_letter_size)])#logにしとく\n\n\nfor i in range(len(id_test)):\n    with Image.open(id_test[i]) as f:\n        width,height=f.size\n        aspect_ratio_pic=height/width\n        aspect_ratio_pic_all_test.append(aspect_ratio_pic)\n\n\nK.clear_session()\nmodel=create_model(input_shape=(input_height,input_width,3),size_detection_mode=True)\n\nlr_schedule = LearningRateScheduler(lrs)\nmodel_checkpoint = ModelCheckpoint(final_weights_step1, monitor = 'val_loss', verbose = 1,\n                                    save_best_only = True, save_weights_only = True, period = 1)\nprint(model.summary())\n\ntrain_list, cv_list = train_test_split(train_input_for_size_estimate, random_state = 111,test_size = 0.2)\n\n\nlearning_rate=0.0005\nn_epoch=10\nbatch_size=32\n\nmodel.compile(loss=mean_squared_error, optimizer=Adam(lr=learning_rate))\nhist = model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size)\n\n#model.save_weights('final_weights_step1.h5')\n#model.load_weights(final_weights_step1)\n</code></pre>\n\n<p>==========================================</p>\n\n<p>Here's the code for the SageMaker notebook file:</p>\n\n<p>import os\nimport keras\nimport numpy as np\nimport sagemaker\nfrom sagemaker.tensorflow import TensorFlow</p>\n\n<p>sess = sagemaker.Session()\nrole = \"arn:aws:iam::xxxxxxxxxxxxxxxxx\"</p>\n\n<p>training_input_path = \"s3://xxxxxxxxxxxx/sagemaker-train-data\"\nvalidation_input_path = \"s3://xxxxxxxxxxxxx/sagemaker-test-data\"</p>\n\n<p>tf_estimator = TensorFlow(entry_point='sagemaker_centernet_script.py', \n                          role=role,\n                          train_instance_count=1, \n                          train_instance_type='ml.p3.2xlarge',\n                          framework_version='1.12', \n                          py_version='py3',\n                          script_mode=True,\n                          hyperparameters={\n                              'epochs': 20,\n                              'batch-size': 256,\n                              'learning-rate': 0.01}\n                         )</p>\n\n<p>tf_estimator.fit({'training': training_input_path, 'validation': validation_input_path}) </p>",
  "messages": [
    {
      "id": "605052",
      "postDate": "08/22/2019 02:33:44",
      "content": "<p>For a long time I have wanted to experiment with <a href=\"https://aws.amazon.com/sagemaker/\">Amazon Sagemaker</a>, so I decided to try it on the notebook with the best score in the Notebooks section of this contest, which is this Centernet Keypoint Detector one from K_mat:  <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">https://www.kaggle.com/kmat2019/centernet-keypoint-detector</a></p>\n\n<p>SageMaker is tool from Amazon that makes using ML easier in various ways, such as being able to run a notebook from your local CPU machine but still do the training part on remote AWS GPU instances. It also has all sorts of built-in ML algorithms that are optimized to be faster than usual, including an optimized version of Tensorflow (which now also includes Keras).</p>\n\n<p>I do not have a GPU myself, so every time I need to use one I first create the notebook as usual on my AWS CPU server, do all sorts of testing, and when I am ready to train the model for real I save an AMI of server (an AMI is like a backup copy) which I then launch a copy of using an AWS GPU server. </p>\n\n<p>But, it would be nice not to have to do that, plus I have to keep checking to see when the GPU server is done with training, or else I would be paying for it while it is no longer in use. So SageMaker seemed like a good possible solution.</p>\n\n<p>I first tried the Centernet Keypoint Detector notebook on my existing AWS CPU server (not using SageMaker) and it got a submission score of 0.515, which is a little less than the original notebook, but because I was using a CPU and not a GPU like he was, it probably trained it less for the same number of epochs. Mainly though, it took around 2 days for training, which is annoyingly long.</p>\n\n<p>I installed SageMaker on my server, and after some time dealing with various configuration issues getting Boto and AWS CLI to work with my AWS account info and S3 buckets, I got SageMaker sort of working using the demo notebooks that came with it.  I am not really a programmer though, and there was one error I could not figure out, so I hired somebody to fix it for me who had experience with SageMaker. He fixed it in a matter of minutes, so I then decided to pay him to get the Centernet Keypoint Detector notebook working for me in SageMaker, because it would have taken my days to do it myself.</p>\n\n<p>It ended up taking him 7 hours, so there is no way I could have done it. He wrote: \"There was some errors for resolving out the dependencies in their training docker, thats where most time was consumed, but it was one time work, Now we have solution for that and can handle that from now on in much less time. It's the total time for building the understanding what is being done in the notebook to running the training job.\"   And also:   \"...issue was related to external libraries like cv2, glob etc. There are workarounds for that to install those, One is implemented and you can use those as placeholders for next projects\".</p>\n\n<p>What I mainly wanted to find out is if using SageMaker (with a GPU) for it was any faster than running the original notebook K_mat posted at:  <a href=\"https://www.kaggle.com/kmat2019/centernet-keypoint-detector\">https://www.kaggle.com/kmat2019/centernet-keypoint-detector</a>.  I am not sure what GPU K-mat used, but with SageMaker I used an AWS V100 which is super fast, so I doubt his was any faster.  For comparison purposes, in step 1 of the training part of his notebook, SageMaker took 801s (9s/step) for the 1st epoch.  K_mat's notebook was much faster at 439s (5s/step).  Most interestingly, my 8 core AWS CPU server (m4.xlarge) took 1844s (20s/step) which was not that much worse, considering how much cheaper it is to use a CPU.</p>\n\n<p>I did not let the SageMaker training finish, so I don't know what the final accuracy would have been, but I have no reason to think it would have been any better than the original notebook. All of this was just about trying to make things either easier or faster for me, but in the end it was not worth it for what I am doing.</p>\n\n<p>Something to keep in mind though is that many people could use SageMaker via the AWS management console (similar to how Kaggle offers Kaggle kernals for you to use), instead of using it locally in script mode. This way, you create the notebook directly in the AWS cloud without running any server yourself.  That might have made it so none of those errors happened, but it still would have involved all the same programming to convert the Kaggle notebook to SageMaker format.</p>\n\n<p>In case anybody is interested, below is all the code I used. SageMaker script works using 2 files. First, it uses the original notebook code as a .py file, with a few simple lines of code added for SageMaker. This .py file is then run as part of a SageMaker notebook, which tells it where your data is stored and which type of AWS instance to use for the training.  </p>\n\n<p>Here's the code:</p>\n\n<p>================================</p>\n\n<p>sagemaker_centernet_script.py</p>\n\n<p>import argparse, os\nos.system(\"apt-get update\")\nos.system(\"apt-get install -y libsm6 libxext6 libxrender-dev\")\nos.system(\"pip install opencv-python-headless\")</p>\n\n<p>os.system(\"pip install Pillow\")\nos.system(\"pip install matplotlib\")\nos.system(\"pip install glob3\")</p>\n\n<p>import numpy as np\nimport json\nimport pandas as pd\nfrom PIL import Image, ImageDraw\nimport matplotlib.pyplot as plt\nfrom pandas.io.json import json_normalize\nimport random\nimport tensorflow as tf\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import KFold,train_test_split\nimport matplotlib.pyplot as plt\nimport glob\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.layers import Dense,Dropout, Conv2D,Conv2DTranspose, BatchNormalization, Activation,AveragePooling2D,GlobalAveragePooling2D, Input, Concatenate, MaxPool2D, Add, UpSampling2D, LeakyReLU,ZeroPadding2D\nfrom keras.models import Model\nfrom keras.objectives import mean_squared_error\nfrom keras import backend as K\nfrom keras.losses import binary_crossentropy\nfrom keras.callbacks import ModelCheckpoint, EarlyStopping, TensorBoard, ReduceLROnPlateau,LearningRateScheduler</p>\n\n<p>from keras.optimizers import Adam, RMSprop, SGD</p>\n\n<p>import boto3 </p>\n\n<p>import keras.backend.tensorflow_backend as K\nK.set_session</p>\n\n<p>category_n=1\nimport cv2\ninput_width,input_height=512, 512</p>\n\n<p>def Datagen_sizecheck_model(filenames, batch_size, size_detection_mode=True, is_train=True,random_crop=True):\n  x=[]\n  y=[]</p>\n\n<p>count=0</p>\n\n<p>while True:\n    for i in range(len(filenames)):\n      if random_crop:\n        crop_ratio=np.random.uniform(0.7,1)\n      else:\n        crop_ratio=1\n      with Image.open(filenames[i][0]) as f:\n        #random crop\n        if random_crop and is_train:\n          pic_width,pic_height=f.size\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8)\n          top_offset=np.random.randint(0,pic_height-int(crop_ratio*pic_height))\n          left_offset=np.random.randint(0,pic_width-int(crop_ratio*pic_width))\n          bottom_offset=top_offset+int(crop_ratio*pic_height)\n          right_offset=left_offset+int(crop_ratio*pic_width)\n          f=cv2.resize(f[top_offset:bottom_offset,left_offset:right_offset,:],(input_height,input_width))\n        else:\n          f=f.resize((input_width, input_height))\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8) <br>\n        x.append(f)</p>\n\n<pre><code>  if random_crop and is_train:\n    y.append(filenames[i][1]-np.log(crop_ratio))\n  else:\n    y.append(filenames[i][1])\n\n  count+=1\n  if count==batch_size:\n    x=np.array(x, dtype=np.float32)\n    y=np.array(y, dtype=np.float32)\n\n    inputs=x/255\n    targets=y       \n    x=[]\n    y=[]\n    count=0\n    yield inputs, targets\n</code></pre>\n\n<p>def aggregation_block(x_shallow, x_deep, deep_ch, out_ch):\n  x_deep= Conv2DTranspose(deep_ch, kernel_size=2, strides=2, padding='same', use_bias=False)(x_deep)\n  x_deep = BatchNormalization()(x_deep) <br>\n  x_deep = LeakyReLU(alpha=0.1)(x_deep)\n  x = Concatenate()([x_shallow, x_deep])\n  x=Conv2D(out_ch, kernel_size=1, strides=1, padding=\"same\")(x)\n  x = BatchNormalization()(x) <br>\n  x = LeakyReLU(alpha=0.1)(x)\n  return x</p>\n\n<p>def cbr(x, out_layer, kernel, stride):\n  x=Conv2D(out_layer, kernel_size=kernel, strides=stride, padding=\"same\")(x)\n  x = BatchNormalization()(x)\n  x = LeakyReLU(alpha=0.1)(x)\n  return x</p>\n\n<p>def resblock(x_in,layer_n):\n  x=cbr(x_in,layer_n,3,1)\n  x=cbr(x,layer_n,3,1)\n  x=Add()([x,x_in])\n  return x  </p>\n\n<h1>I use the same network at CenterNet</h1>\n\n<p>def create_model(input_shape, size_detection_mode=True, aggregation=True):\n    input_layer = Input(input_shape)</p>\n\n<pre><code>#resized input\ninput_layer_1=AveragePooling2D(2)(input_layer)\ninput_layer_2=AveragePooling2D(2)(input_layer_1)\n\n#### ENCODER ####\n\nx_0= cbr(input_layer, 16, 3, 2)#512-&gt;256\nconcat_1 = Concatenate()([x_0, input_layer_1])\n\nx_1= cbr(concat_1, 32, 3, 2)#256-&gt;128\nconcat_2 = Concatenate()([x_1, input_layer_2])\n\nx_2= cbr(concat_2, 64, 3, 2)#128-&gt;64\n\nx=cbr(x_2,64,3,1)\nx=resblock(x,64)\nx=resblock(x,64)\n\nx_3= cbr(x, 128, 3, 2)#64-&gt;32\nx= cbr(x_3, 128, 3, 1)\nx=resblock(x,128)\nx=resblock(x,128)\nx=resblock(x,128)\n\nx_4= cbr(x, 256, 3, 2)#32-&gt;16\nx= cbr(x_4, 256, 3, 1)\nx=resblock(x,256)\nx=resblock(x,256)\nx=resblock(x,256)\nx=resblock(x,256)\nx=resblock(x,256)\n\nx_5= cbr(x, 512, 3, 2)#16-&gt;8\nx= cbr(x_5, 512, 3, 1)\n\nx=resblock(x,512)\nx=resblock(x,512)\nx=resblock(x,512)\n\nif size_detection_mode:\n  x=GlobalAveragePooling2D()(x)\n  x=Dropout(0.2)(x)\n  out=Dense(1,activation=\"linear\")(x)\n\nelse:#centernet mode\n#### DECODER ####\n  x_1= cbr(x_1, output_layer_n, 1, 1)\n  x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n  x_2= cbr(x_2, output_layer_n, 1, 1)\n  x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n  x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n  x_3= cbr(x_3, output_layer_n, 1, 1)\n  x_3 = aggregation_block(x_3, x_4, output_layer_n, output_layer_n) \n  x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n  x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n\n  x_4= cbr(x_4, output_layer_n, 1, 1)\n\n  x=cbr(x, output_layer_n, 1, 1)\n  x= UpSampling2D(size=(2, 2))(x)#8-&gt;16 tconvのがいいか\n\n  x = Concatenate()([x, x_4])\n  x=cbr(x, output_layer_n, 3, 1)\n  x= UpSampling2D(size=(2, 2))(x)#16-&gt;32\n\n  x = Concatenate()([x, x_3])\n  x=cbr(x, output_layer_n, 3, 1)\n  x= UpSampling2D(size=(2, 2))(x)#32-&gt;64   128のがいいかも？ \n\n  x = Concatenate()([x, x_2])\n  x=cbr(x, output_layer_n, 3, 1)\n  x= UpSampling2D(size=(2, 2))(x)#64-&gt;128 \n\n  x = Concatenate()([x, x_1])\n  x=Conv2D(output_layer_n, kernel_size=3, strides=1, padding=\"same\")(x)\n  out = Activation(\"sigmoid\")(x)\n\nmodel=Model(input_layer, out)\n\nreturn model\n</code></pre>\n\n<p>def model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size=32):\n    hist = model.fit_generator(\n        Datagen_sizecheck_model(train_list,batch_size, is_train=True,random_crop=True),\n        steps_per_epoch = len(train_list) // batch_size,\n        epochs = n_epoch,\n        validation_data=Datagen_sizecheck_model(cv_list,batch_size, is_train=False,random_crop=False),\n        validation_steps = len(cv_list) // batch_size,\n        callbacks = [lr_schedule, model_checkpoint],#[early_stopping, reduce_lr, model_checkpoint],\n        shuffle = True,\n        verbose = 1\n    )\n    return hist</p>\n\n<p>def lrs(epoch):\n    lr = 0.0005\n    if epoch&gt;10:\n        lr = 0.0001\n    return lr</p>\n\n<p>if <strong>name</strong> == '<strong>main</strong>':</p>\n\n<pre><code>parser = argparse.ArgumentParser()\n\nparser.add_argument('--epochs', type=int, default=10)\nparser.add_argument('--learning-rate', type=float, default=0.01)\nparser.add_argument('--batch-size', type=int, default=128)\nparser.add_argument('--gpu-count', type=int, default=os.environ['SM_NUM_GPUS'])\nparser.add_argument('--model-dir', type=str, default=os.environ['SM_MODEL_DIR'])\nparser.add_argument('--training', type=str, default=os.environ['SM_CHANNEL_TRAINING'])\nparser.add_argument('--validation', type=str, default=os.environ['SM_CHANNEL_VALIDATION'])\n\nargs, _ = parser.parse_known_args()\ns3 = boto3.client('s3')\n\nepochs     = args.epochs\nlr         = args.learning_rate\nbatch_size = args.batch_size\ngpu_count  = args.gpu_count\nmodel_dir  = args.model_dir\ntraining_dir   = args.training\nvalidation_dir = args.validation\n\npath_1= \"/opt/ml/input/data/training/train_csv/train.csv\"\npath_2=\"/opt/ml/input/data/training/\"\npath_3=\"/opt/ml/input/data/validation/\"\npath_4=\"/opt/ml/input/data/validation/submission/sample_submission.csv\"\nfinal_weights_step1 = \"/opt/ml/input/data/training/final_weights/final_weights_step1.hdf5\"\n\n\ndf_train=pd.read_csv(path_1)\n#print(df_train.head())\n#print(df_train.shape)\ndf_train=df_train.dropna(axis=0, how='any')#you can use nan data(page with no letter)\ndf_train=df_train.reset_index(drop=True)\n#print(df_train.shape)\n\nannotation_list_train=[]\ncategory_names=set()\n\nfor i in range(len(df_train)):\n    ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,x,y,width,height for each picture\n    category_names=category_names.union({i for i in ann[:,0]})\n\ncategory_names=sorted(category_names)\ndict_cat={list(category_names)[j]:str(j) for j in range(len(category_names))}\ninv_dict_cat={str(j):list(category_names)[j] for j in range(len(category_names))}\n#print(dict_cat)\n\nfor i in range(len(df_train)):\n    ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,left,top,width,height for each picture\n    for j,category_name in enumerate(ann[:,0]):\n        ann[j,0]=int(dict_cat[category_name])  \n    ann=ann.astype('int32')\n    ann[:,1]+=ann[:,3]//2#center_x\n    ann[:,2]+=ann[:,4]//2#center_y\n    annotation_list_train.append([\"{}{}.jpg\".format(path_2,df_train.loc[i,\"image_id\"]),ann])\n\ndf_submission=pd.read_csv(path_4)\nid_test=path_3+df_submission[\"image_id\"]+\".jpg\"\n\naspect_ratio_pic_all=[]\naspect_ratio_pic_all_test=[]\naverage_letter_size_all=[]\ntrain_input_for_size_estimate=[]\nresize_dir=\"resized/\"\nif os.path.exists(resize_dir) == False:os.mkdir(resize_dir)\nfor i in range(len(annotation_list_train)):\n    with Image.open(annotation_list_train[i][0]) as f:\n        width,height=f.size\n        area=width*height\n        aspect_ratio_pic=height/width\n        aspect_ratio_pic_all.append(aspect_ratio_pic)\n        letter_size=annotation_list_train[i][1][:,3]*annotation_list_train[i][1][:,4]\n        letter_size_ratio=letter_size/area\n\n        average_letter_size=np.mean(letter_size_ratio)\n        average_letter_size_all.append(average_letter_size)\n        train_input_for_size_estimate.append([annotation_list_train[i][0],np.log(average_letter_size)])#logにしとく\n\n\nfor i in range(len(id_test)):\n    with Image.open(id_test[i]) as f:\n        width,height=f.size\n        aspect_ratio_pic=height/width\n        aspect_ratio_pic_all_test.append(aspect_ratio_pic)\n\n\nK.clear_session()\nmodel=create_model(input_shape=(input_height,input_width,3),size_detection_mode=True)\n\nlr_schedule = LearningRateScheduler(lrs)\nmodel_checkpoint = ModelCheckpoint(final_weights_step1, monitor = 'val_loss', verbose = 1,\n                                    save_best_only = True, save_weights_only = True, period = 1)\nprint(model.summary())\n\ntrain_list, cv_list = train_test_split(train_input_for_size_estimate, random_state = 111,test_size = 0.2)\n\n\nlearning_rate=0.0005\nn_epoch=10\nbatch_size=32\n\nmodel.compile(loss=mean_squared_error, optimizer=Adam(lr=learning_rate))\nhist = model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size)\n\n#model.save_weights('final_weights_step1.h5')\n#model.load_weights(final_weights_step1)\n</code></pre>\n\n<p>==========================================</p>\n\n<p>Here's the code for the SageMaker notebook file:</p>\n\n<p>import os\nimport keras\nimport numpy as np\nimport sagemaker\nfrom sagemaker.tensorflow import TensorFlow</p>\n\n<p>sess = sagemaker.Session()\nrole = \"arn:aws:iam::xxxxxxxxxxxxxxxxx\"</p>\n\n<p>training_input_path = \"s3://xxxxxxxxxxxx/sagemaker-train-data\"\nvalidation_input_path = \"s3://xxxxxxxxxxxxx/sagemaker-test-data\"</p>\n\n<p>tf_estimator = TensorFlow(entry_point='sagemaker_centernet_script.py', \n                          role=role,\n                          train_instance_count=1, \n                          train_instance_type='ml.p3.2xlarge',\n                          framework_version='1.12', \n                          py_version='py3',\n                          script_mode=True,\n                          hyperparameters={\n                              'epochs': 20,\n                              'batch-size': 256,\n                              'learning-rate': 0.01}\n                         )</p>\n\n<p>tf_estimator.fit({'training': training_input_path, 'validation': validation_input_path}) </p>",
      "rawMarkdown": "For a long time I have wanted to experiment with [Amazon Sagemaker](https://aws.amazon.com/sagemaker/), so I decided to try it on the notebook with the best score in the Notebooks section of this contest, which is this Centernet Keypoint Detector one from K_mat:  [https://www.kaggle.com/kmat2019/centernet-keypoint-detector](https://www.kaggle.com/kmat2019/centernet-keypoint-detector)\n\nSageMaker is tool from Amazon that makes using ML easier in various ways, such as being able to run a notebook from your local CPU machine but still do the training part on remote AWS GPU instances. It also has all sorts of built-in ML algorithms that are optimized to be faster than usual, including an optimized version of Tensorflow (which now also includes Keras).\n\nI do not have a GPU myself, so every time I need to use one I first create the notebook as usual on my AWS CPU server, do all sorts of testing, and when I am ready to train the model for real I save an AMI of server (an AMI is like a backup copy) which I then launch a copy of using an AWS GPU server. \n\nBut, it would be nice not to have to do that, plus I have to keep checking to see when the GPU server is done with training, or else I would be paying for it while it is no longer in use. So SageMaker seemed like a good possible solution.\n\nI first tried the Centernet Keypoint Detector notebook on my existing AWS CPU server (not using SageMaker) and it got a submission score of 0.515, which is a little less than the original notebook, but because I was using a CPU and not a GPU like he was, it probably trained it less for the same number of epochs. Mainly though, it took around 2 days for training, which is annoyingly long.\n\nI installed SageMaker on my server, and after some time dealing with various configuration issues getting Boto and AWS CLI to work with my AWS account info and S3 buckets, I got SageMaker sort of working using the demo notebooks that came with it.  I am not really a programmer though, and there was one error I could not figure out, so I hired somebody to fix it for me who had experience with SageMaker. He fixed it in a matter of minutes, so I then decided to pay him to get the Centernet Keypoint Detector notebook working for me in SageMaker, because it would have taken my days to do it myself.\n\nIt ended up taking him 7 hours, so there is no way I could have done it. He wrote: \"There was some errors for resolving out the dependencies in their training docker, thats where most time was consumed, but it was one time work, Now we have solution for that and can handle that from now on in much less time. It's the total time for building the understanding what is being done in the notebook to running the training job.\"   And also:   \"...issue was related to external libraries like cv2, glob etc. There are workarounds for that to install those, One is implemented and you can use those as placeholders for next projects\".\n\nWhat I mainly wanted to find out is if using SageMaker (with a GPU) for it was any faster than running the original notebook K_mat posted at:  [https://www.kaggle.com/kmat2019/centernet-keypoint-detector](https://www.kaggle.com/kmat2019/centernet-keypoint-detector).  I am not sure what GPU K-mat used, but with SageMaker I used an AWS V100 which is super fast, so I doubt his was any faster.  For comparison purposes, in step 1 of the training part of his notebook, SageMaker took 801s (9s/step) for the 1st epoch.  K_mat's notebook was much faster at 439s (5s/step).  Most interestingly, my 8 core AWS CPU server (m4.xlarge) took 1844s (20s/step) which was not that much worse, considering how much cheaper it is to use a CPU.\n\nI did not let the SageMaker training finish, so I don't know what the final accuracy would have been, but I have no reason to think it would have been any better than the original notebook. All of this was just about trying to make things either easier or faster for me, but in the end it was not worth it for what I am doing.\n\nSomething to keep in mind though is that many people could use SageMaker via the AWS management console (similar to how Kaggle offers Kaggle kernals for you to use), instead of using it locally in script mode. This way, you create the notebook directly in the AWS cloud without running any server yourself.  That might have made it so none of those errors happened, but it still would have involved all the same programming to convert the Kaggle notebook to SageMaker format.\n\nIn case anybody is interested, below is all the code I used. SageMaker script works using 2 files. First, it uses the original notebook code as a .py file, with a few simple lines of code added for SageMaker. This .py file is then run as part of a SageMaker notebook, which tells it where your data is stored and which type of AWS instance to use for the training.  \n\nHere's the code:\n\n================================\n\nsagemaker_centernet_script.py\n\nimport argparse, os\nos.system(\"apt-get update\")\nos.system(\"apt-get install -y libsm6 libxext6 libxrender-dev\")\nos.system(\"pip install opencv-python-headless\")\n\nos.system(\"pip install Pillow\")\nos.system(\"pip install matplotlib\")\nos.system(\"pip install glob3\")\n\nimport numpy as np\nimport json\nimport pandas as pd\nfrom PIL import Image, ImageDraw\nimport matplotlib.pyplot as plt\nfrom pandas.io.json import json_normalize\nimport random\nimport tensorflow as tf\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import KFold,train_test_split\nimport matplotlib.pyplot as plt\nimport glob\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.layers import Dense,Dropout, Conv2D,Conv2DTranspose, BatchNormalization, Activation,AveragePooling2D,GlobalAveragePooling2D, Input, Concatenate, MaxPool2D, Add, UpSampling2D, LeakyReLU,ZeroPadding2D\nfrom keras.models import Model\nfrom keras.objectives import mean_squared_error\nfrom keras import backend as K\nfrom keras.losses import binary_crossentropy\nfrom keras.callbacks import ModelCheckpoint, EarlyStopping, TensorBoard, ReduceLROnPlateau,LearningRateScheduler\n  \nfrom keras.optimizers import Adam, RMSprop, SGD\n\nimport boto3 \n\nimport keras.backend.tensorflow_backend as K\nK.set_session\n\ncategory_n=1\nimport cv2\ninput_width,input_height=512, 512\n\ndef Datagen_sizecheck_model(filenames, batch_size, size_detection_mode=True, is_train=True,random_crop=True):\n  x=[]\n  y=[]\n  \n  count=0\n\n  while True:\n    for i in range(len(filenames)):\n      if random_crop:\n        crop_ratio=np.random.uniform(0.7,1)\n      else:\n        crop_ratio=1\n      with Image.open(filenames[i][0]) as f:\n        #random crop\n        if random_crop and is_train:\n          pic_width,pic_height=f.size\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8)\n          top_offset=np.random.randint(0,pic_height-int(crop_ratio*pic_height))\n          left_offset=np.random.randint(0,pic_width-int(crop_ratio*pic_width))\n          bottom_offset=top_offset+int(crop_ratio*pic_height)\n          right_offset=left_offset+int(crop_ratio*pic_width)\n          f=cv2.resize(f[top_offset:bottom_offset,left_offset:right_offset,:],(input_height,input_width))\n        else:\n          f=f.resize((input_width, input_height))\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8)          \n        x.append(f)\n      \n      \n      if random_crop and is_train:\n        y.append(filenames[i][1]-np.log(crop_ratio))\n      else:\n        y.append(filenames[i][1])\n      \n      count+=1\n      if count==batch_size:\n        x=np.array(x, dtype=np.float32)\n        y=np.array(y, dtype=np.float32)\n\n        inputs=x/255\n        targets=y       \n        x=[]\n        y=[]\n        count=0\n        yield inputs, targets\n\n\n\ndef aggregation_block(x_shallow, x_deep, deep_ch, out_ch):\n  x_deep= Conv2DTranspose(deep_ch, kernel_size=2, strides=2, padding='same', use_bias=False)(x_deep)\n  x_deep = BatchNormalization()(x_deep)   \n  x_deep = LeakyReLU(alpha=0.1)(x_deep)\n  x = Concatenate()([x_shallow, x_deep])\n  x=Conv2D(out_ch, kernel_size=1, strides=1, padding=\"same\")(x)\n  x = BatchNormalization()(x)   \n  x = LeakyReLU(alpha=0.1)(x)\n  return x\n  \n\n\ndef cbr(x, out_layer, kernel, stride):\n  x=Conv2D(out_layer, kernel_size=kernel, strides=stride, padding=\"same\")(x)\n  x = BatchNormalization()(x)\n  x = LeakyReLU(alpha=0.1)(x)\n  return x\n\ndef resblock(x_in,layer_n):\n  x=cbr(x_in,layer_n,3,1)\n  x=cbr(x,layer_n,3,1)\n  x=Add()([x,x_in])\n  return x  \n\n\n#I use the same network at CenterNet\ndef create_model(input_shape, size_detection_mode=True, aggregation=True):\n    input_layer = Input(input_shape)\n    \n    #resized input\n    input_layer_1=AveragePooling2D(2)(input_layer)\n    input_layer_2=AveragePooling2D(2)(input_layer_1)\n\n    #### ENCODER ####\n\n    x_0= cbr(input_layer, 16, 3, 2)#512-&gt;256\n    concat_1 = Concatenate()([x_0, input_layer_1])\n\n    x_1= cbr(concat_1, 32, 3, 2)#256-&gt;128\n    concat_2 = Concatenate()([x_1, input_layer_2])\n\n    x_2= cbr(concat_2, 64, 3, 2)#128-&gt;64\n    \n    x=cbr(x_2,64,3,1)\n    x=resblock(x,64)\n    x=resblock(x,64)\n    \n    x_3= cbr(x, 128, 3, 2)#64-&gt;32\n    x= cbr(x_3, 128, 3, 1)\n    x=resblock(x,128)\n    x=resblock(x,128)\n    x=resblock(x,128)\n    \n    x_4= cbr(x, 256, 3, 2)#32-&gt;16\n    x= cbr(x_4, 256, 3, 1)\n    x=resblock(x,256)\n    x=resblock(x,256)\n    x=resblock(x,256)\n    x=resblock(x,256)\n    x=resblock(x,256)\n \n    x_5= cbr(x, 512, 3, 2)#16-&gt;8\n    x= cbr(x_5, 512, 3, 1)\n    \n    x=resblock(x,512)\n    x=resblock(x,512)\n    x=resblock(x,512)\n    \n    if size_detection_mode:\n      x=GlobalAveragePooling2D()(x)\n      x=Dropout(0.2)(x)\n      out=Dense(1,activation=\"linear\")(x)\n    \n    else:#centernet mode\n    #### DECODER ####\n      x_1= cbr(x_1, output_layer_n, 1, 1)\n      x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n      x_2= cbr(x_2, output_layer_n, 1, 1)\n      x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n      x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n      x_3= cbr(x_3, output_layer_n, 1, 1)\n      x_3 = aggregation_block(x_3, x_4, output_layer_n, output_layer_n) \n      x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n      x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n      \n      x_4= cbr(x_4, output_layer_n, 1, 1)\n\n      x=cbr(x, output_layer_n, 1, 1)\n      x= UpSampling2D(size=(2, 2))(x)#8-&gt;16 tconvのがいいか\n\n      x = Concatenate()([x, x_4])\n      x=cbr(x, output_layer_n, 3, 1)\n      x= UpSampling2D(size=(2, 2))(x)#16-&gt;32\n    \n      x = Concatenate()([x, x_3])\n      x=cbr(x, output_layer_n, 3, 1)\n      x= UpSampling2D(size=(2, 2))(x)#32-&gt;64   128のがいいかも？ \n    \n      x = Concatenate()([x, x_2])\n      x=cbr(x, output_layer_n, 3, 1)\n      x= UpSampling2D(size=(2, 2))(x)#64-&gt;128 \n      \n      x = Concatenate()([x, x_1])\n      x=Conv2D(output_layer_n, kernel_size=3, strides=1, padding=\"same\")(x)\n      out = Activation(\"sigmoid\")(x)\n    \n    model=Model(input_layer, out)\n    \n    return model\n  \n    \n\ndef model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size=32):\n    hist = model.fit_generator(\n        Datagen_sizecheck_model(train_list,batch_size, is_train=True,random_crop=True),\n        steps_per_epoch = len(train_list) // batch_size,\n        epochs = n_epoch,\n        validation_data=Datagen_sizecheck_model(cv_list,batch_size, is_train=False,random_crop=False),\n        validation_steps = len(cv_list) // batch_size,\n        callbacks = [lr_schedule, model_checkpoint],#[early_stopping, reduce_lr, model_checkpoint],\n        shuffle = True,\n        verbose = 1\n    )\n    return hist\n\ndef lrs(epoch):\n    lr = 0.0005\n    if epoch&gt;10:\n        lr = 0.0001\n    return lr\n\nif __name__ == '__main__':\n        \n    parser = argparse.ArgumentParser()\n\n    parser.add_argument('--epochs', type=int, default=10)\n    parser.add_argument('--learning-rate', type=float, default=0.01)\n    parser.add_argument('--batch-size', type=int, default=128)\n    parser.add_argument('--gpu-count', type=int, default=os.environ['SM_NUM_GPUS'])\n    parser.add_argument('--model-dir', type=str, default=os.environ['SM_MODEL_DIR'])\n    parser.add_argument('--training', type=str, default=os.environ['SM_CHANNEL_TRAINING'])\n    parser.add_argument('--validation', type=str, default=os.environ['SM_CHANNEL_VALIDATION'])\n    \n    args, _ = parser.parse_known_args()\n    s3 = boto3.client('s3')\n\n    epochs     = args.epochs\n    lr         = args.learning_rate\n    batch_size = args.batch_size\n    gpu_count  = args.gpu_count\n    model_dir  = args.model_dir\n    training_dir   = args.training\n    validation_dir = args.validation\n\n    path_1= \"/opt/ml/input/data/training/train_csv/train.csv\"\n    path_2=\"/opt/ml/input/data/training/\"\n    path_3=\"/opt/ml/input/data/validation/\"\n    path_4=\"/opt/ml/input/data/validation/submission/sample_submission.csv\"\n    final_weights_step1 = \"/opt/ml/input/data/training/final_weights/final_weights_step1.hdf5\"\n\n\n    df_train=pd.read_csv(path_1)\n    #print(df_train.head())\n    #print(df_train.shape)\n    df_train=df_train.dropna(axis=0, how='any')#you can use nan data(page with no letter)\n    df_train=df_train.reset_index(drop=True)\n    #print(df_train.shape)\n\n    annotation_list_train=[]\n    category_names=set()\n\n    for i in range(len(df_train)):\n        ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,x,y,width,height for each picture\n        category_names=category_names.union({i for i in ann[:,0]})\n\n    category_names=sorted(category_names)\n    dict_cat={list(category_names)[j]:str(j) for j in range(len(category_names))}\n    inv_dict_cat={str(j):list(category_names)[j] for j in range(len(category_names))}\n    #print(dict_cat)\n    \n    for i in range(len(df_train)):\n        ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,left,top,width,height for each picture\n        for j,category_name in enumerate(ann[:,0]):\n            ann[j,0]=int(dict_cat[category_name])  \n        ann=ann.astype('int32')\n        ann[:,1]+=ann[:,3]//2#center_x\n        ann[:,2]+=ann[:,4]//2#center_y\n        annotation_list_train.append([\"{}{}.jpg\".format(path_2,df_train.loc[i,\"image_id\"]),ann])\n\n    df_submission=pd.read_csv(path_4)\n    id_test=path_3+df_submission[\"image_id\"]+\".jpg\"\n\n    aspect_ratio_pic_all=[]\n    aspect_ratio_pic_all_test=[]\n    average_letter_size_all=[]\n    train_input_for_size_estimate=[]\n    resize_dir=\"resized/\"\n    if os.path.exists(resize_dir) == False:os.mkdir(resize_dir)\n    for i in range(len(annotation_list_train)):\n        with Image.open(annotation_list_train[i][0]) as f:\n            width,height=f.size\n            area=width*height\n            aspect_ratio_pic=height/width\n            aspect_ratio_pic_all.append(aspect_ratio_pic)\n            letter_size=annotation_list_train[i][1][:,3]*annotation_list_train[i][1][:,4]\n            letter_size_ratio=letter_size/area\n        \n            average_letter_size=np.mean(letter_size_ratio)\n            average_letter_size_all.append(average_letter_size)\n            train_input_for_size_estimate.append([annotation_list_train[i][0],np.log(average_letter_size)])#logにしとく\n        \n\n    for i in range(len(id_test)):\n        with Image.open(id_test[i]) as f:\n            width,height=f.size\n            aspect_ratio_pic=height/width\n            aspect_ratio_pic_all_test.append(aspect_ratio_pic)\n\n    \n    K.clear_session()\n    model=create_model(input_shape=(input_height,input_width,3),size_detection_mode=True)\n\n    lr_schedule = LearningRateScheduler(lrs)\n    model_checkpoint = ModelCheckpoint(final_weights_step1, monitor = 'val_loss', verbose = 1,\n                                        save_best_only = True, save_weights_only = True, period = 1)\n    print(model.summary())\n\n    train_list, cv_list = train_test_split(train_input_for_size_estimate, random_state = 111,test_size = 0.2)\n\n\n    learning_rate=0.0005\n    n_epoch=10\n    batch_size=32\n\n    model.compile(loss=mean_squared_error, optimizer=Adam(lr=learning_rate))\n    hist = model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size)\n\n    #model.save_weights('final_weights_step1.h5')\n    #model.load_weights(final_weights_step1)\n\n==========================================\n\nHere's the code for the SageMaker notebook file:\n\nimport os\nimport keras\nimport numpy as np\nimport sagemaker\nfrom sagemaker.tensorflow import TensorFlow\n\nsess = sagemaker.Session()\nrole = \"arn:aws:iam::xxxxxxxxxxxxxxxxx\"\n\ntraining_input_path = \"s3://xxxxxxxxxxxx/sagemaker-train-data\"\nvalidation_input_path = \"s3://xxxxxxxxxxxxx/sagemaker-test-data\"\n\ntf_estimator = TensorFlow(entry_point='sagemaker_centernet_script.py', \n                          role=role,\n                          train_instance_count=1, \n                          train_instance_type='ml.p3.2xlarge',\n                          framework_version='1.12', \n                          py_version='py3',\n                          script_mode=True,\n                          hyperparameters={\n                              'epochs': 20,\n                              'batch-size': 256,\n                              'learning-rate': 0.01}\n                         )\n\ntf_estimator.fit({'training': training_input_path, 'validation': validation_input_path})",
      "votes": null
    },
    {
      "id": "605068",
      "postDate": "08/22/2019 02:57:22",
      "content": "<p>The slowdown is mostly explainable to the data generators not using keras.utils.Sequence, which when used enables multithreaded batch generation.  With the right code, I get avg 138 s/epoch at 2 s/step with workers=8 and max_queue_size=16. Note with Juypter and python use_multiprocessing should always be false due to a bug (unless you do some sleuthing to find some code that properly circumvents the known issue).  Additional note is that Step 1 cannot be too accurate as I've found that it breaks the heatmap into detecting everything on a page.</p>",
      "rawMarkdown": "The slowdown is mostly explainable to the data generators not using keras.utils.Sequence, which when used enables multithreaded batch generation.  With the right code, I get avg 138 s/epoch at 2 s/step with workers=8 and max\\_queue\\_size=16. Note with Juypter and python use\\_multiprocessing should always be false due to a bug (unless you do some sleuthing to find some code that properly circumvents the known issue).  Additional note is that Step 1 cannot be too accurate as I've found that it breaks the heatmap into detecting everything on a page.",
      "votes": null
    },
    {
      "id": "605071",
      "postDate": "08/22/2019 03:00:27",
      "content": "<p>Thanks, that is very interesting. But I was using the same exact settings in my CPU notebook as with my SageMaker notebook, which were also both the same as the original Kaggle notebook. So for comparison purposes like I was doing, would those settings you mentioned have mattered? </p>",
      "rawMarkdown": "Thanks, that is very interesting. But I was using the same exact settings in my CPU notebook as with my SageMaker notebook, which were also both the same as the original Kaggle notebook. So for comparison purposes like I was doing, would those settings you mentioned have mattered?",
      "votes": null
    },
    {
      "id": "605076",
      "postDate": "08/22/2019 03:11:52",
      "content": "<p>So far from my experience, yes it would as setting up multithreading and having a good balance of threads/workers to gpu compute speed is shaving hours off on training but I am not sure on CPU only.  I am using my personal desktop with a Threadripper 1900X and a 1080 Ti for my numbers.  Surely with the SageMaker gpu instance you've used there is an extreme benefit.  I currently cannot test a CPU only config, but I think there is at least minimal benefit.</p>",
      "rawMarkdown": "So far from my experience, yes it would as setting up multithreading and having a good balance of threads/workers to gpu compute speed is shaving hours off on training but I am not sure on CPU only.  I am using my personal desktop with a Threadripper 1900X and a 1080 Ti for my numbers.  Surely with the SageMaker gpu instance you've used there is an extreme benefit.  I currently cannot test a CPU only config, but I think there is at least minimal benefit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 605068,
      "author_name": "cryptosnail",
      "author_url": "",
      "post_date": "08/22/2019 02:57:22",
      "content": "<p>The slowdown is mostly explainable to the data generators not using keras.utils.Sequence, which when used enables multithreaded batch generation.  With the right code, I get avg 138 s/epoch at 2 s/step with workers=8 and max_queue_size=16. Note with Juypter and python use_multiprocessing should always be false due to a bug (unless you do some sleuthing to find some code that properly circumvents the known issue).  Additional note is that Step 1 cannot be too accurate as I've found that it breaks the heatmap into detecting everything on a page.</p>",
      "votes": null,
      "replies": [
        {
          "id": 605071,
          "author_name": "impulsecorp",
          "author_url": "",
          "post_date": "08/22/2019 03:00:27",
          "content": "<p>Thanks, that is very interesting. But I was using the same exact settings in my CPU notebook as with my SageMaker notebook, which were also both the same as the original Kaggle notebook. So for comparison purposes like I was doing, would those settings you mentioned have mattered? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 605076,
          "author_name": "cryptosnail",
          "author_url": "",
          "post_date": "08/22/2019 03:11:52",
          "content": "<p>So far from my experience, yes it would as setting up multithreading and having a good balance of threads/workers to gpu compute speed is shaving hours off on training but I am not sure on CPU only.  I am using my personal desktop with a Threadripper 1900X and a 1080 Ti for my numbers.  Surely with the SageMaker gpu instance you've used there is an extreme benefit.  I currently cannot test a CPU only config, but I think there is at least minimal benefit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "605052": "For a long time I have wanted to experiment with [Amazon Sagemaker](https://aws.amazon.com/sagemaker/), so I decided to try it on the notebook with the best score in the Notebooks section of this contest, which is this Centernet Keypoint Detector one from K_mat:  [https://www.kaggle.com/kmat2019/centernet-keypoint-detector](https://www.kaggle.com/kmat2019/centernet-keypoint-detector)\n\nSageMaker is tool from Amazon that makes using ML easier in various ways, such as being able to run a notebook from your local CPU machine but still do the training part on remote AWS GPU instances. It also has all sorts of built-in ML algorithms that are optimized to be faster than usual, including an optimized version of Tensorflow (which now also includes Keras).\n\nI do not have a GPU myself, so every time I need to use one I first create the notebook as usual on my AWS CPU server, do all sorts of testing, and when I am ready to train the model for real I save an AMI of server (an AMI is like a backup copy) which I then launch a copy of using an AWS GPU server. \n\nBut, it would be nice not to have to do that, plus I have to keep checking to see when the GPU server is done with training, or else I would be paying for it while it is no longer in use. So SageMaker seemed like a good possible solution.\n\nI first tried the Centernet Keypoint Detector notebook on my existing AWS CPU server (not using SageMaker) and it got a submission score of 0.515, which is a little less than the original notebook, but because I was using a CPU and not a GPU like he was, it probably trained it less for the same number of epochs. Mainly though, it took around 2 days for training, which is annoyingly long.\n\nI installed SageMaker on my server, and after some time dealing with various configuration issues getting Boto and AWS CLI to work with my AWS account info and S3 buckets, I got SageMaker sort of working using the demo notebooks that came with it.  I am not really a programmer though, and there was one error I could not figure out, so I hired somebody to fix it for me who had experience with SageMaker. He fixed it in a matter of minutes, so I then decided to pay him to get the Centernet Keypoint Detector notebook working for me in SageMaker, because it would have taken my days to do it myself.\n\nIt ended up taking him 7 hours, so there is no way I could have done it. He wrote: \"There was some errors for resolving out the dependencies in their training docker, thats where most time was consumed, but it was one time work, Now we have solution for that and can handle that from now on in much less time. It's the total time for building the understanding what is being done in the notebook to running the training job.\"   And also:   \"...issue was related to external libraries like cv2, glob etc. There are workarounds for that to install those, One is implemented and you can use those as placeholders for next projects\".\n\nWhat I mainly wanted to find out is if using SageMaker (with a GPU) for it was any faster than running the original notebook K_mat posted at:  [https://www.kaggle.com/kmat2019/centernet-keypoint-detector](https://www.kaggle.com/kmat2019/centernet-keypoint-detector).  I am not sure what GPU K-mat used, but with SageMaker I used an AWS V100 which is super fast, so I doubt his was any faster.  For comparison purposes, in step 1 of the training part of his notebook, SageMaker took 801s (9s/step) for the 1st epoch.  K_mat's notebook was much faster at 439s (5s/step).  Most interestingly, my 8 core AWS CPU server (m4.xlarge) took 1844s (20s/step) which was not that much worse, considering how much cheaper it is to use a CPU.\n\nI did not let the SageMaker training finish, so I don't know what the final accuracy would have been, but I have no reason to think it would have been any better than the original notebook. All of this was just about trying to make things either easier or faster for me, but in the end it was not worth it for what I am doing.\n\nSomething to keep in mind though is that many people could use SageMaker via the AWS management console (similar to how Kaggle offers Kaggle kernals for you to use), instead of using it locally in script mode. This way, you create the notebook directly in the AWS cloud without running any server yourself.  That might have made it so none of those errors happened, but it still would have involved all the same programming to convert the Kaggle notebook to SageMaker format.\n\nIn case anybody is interested, below is all the code I used. SageMaker script works using 2 files. First, it uses the original notebook code as a .py file, with a few simple lines of code added for SageMaker. This .py file is then run as part of a SageMaker notebook, which tells it where your data is stored and which type of AWS instance to use for the training.  \n\nHere's the code:\n\n================================\n\nsagemaker_centernet_script.py\n\nimport argparse, os\nos.system(\"apt-get update\")\nos.system(\"apt-get install -y libsm6 libxext6 libxrender-dev\")\nos.system(\"pip install opencv-python-headless\")\n\nos.system(\"pip install Pillow\")\nos.system(\"pip install matplotlib\")\nos.system(\"pip install glob3\")\n\nimport numpy as np\nimport json\nimport pandas as pd\nfrom PIL import Image, ImageDraw\nimport matplotlib.pyplot as plt\nfrom pandas.io.json import json_normalize\nimport random\nimport tensorflow as tf\nfrom sklearn.utils import shuffle\nfrom sklearn.model_selection import KFold,train_test_split\nimport matplotlib.pyplot as plt\nimport glob\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.layers import Dense,Dropout, Conv2D,Conv2DTranspose, BatchNormalization, Activation,AveragePooling2D,GlobalAveragePooling2D, Input, Concatenate, MaxPool2D, Add, UpSampling2D, LeakyReLU,ZeroPadding2D\nfrom keras.models import Model\nfrom keras.objectives import mean_squared_error\nfrom keras import backend as K\nfrom keras.losses import binary_crossentropy\nfrom keras.callbacks import ModelCheckpoint, EarlyStopping, TensorBoard, ReduceLROnPlateau,LearningRateScheduler\n  \nfrom keras.optimizers import Adam, RMSprop, SGD\n\nimport boto3 \n\nimport keras.backend.tensorflow_backend as K\nK.set_session\n\ncategory_n=1\nimport cv2\ninput_width,input_height=512, 512\n\ndef Datagen_sizecheck_model(filenames, batch_size, size_detection_mode=True, is_train=True,random_crop=True):\n  x=[]\n  y=[]\n  \n  count=0\n\n  while True:\n    for i in range(len(filenames)):\n      if random_crop:\n        crop_ratio=np.random.uniform(0.7,1)\n      else:\n        crop_ratio=1\n      with Image.open(filenames[i][0]) as f:\n        #random crop\n        if random_crop and is_train:\n          pic_width,pic_height=f.size\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8)\n          top_offset=np.random.randint(0,pic_height-int(crop_ratio*pic_height))\n          left_offset=np.random.randint(0,pic_width-int(crop_ratio*pic_width))\n          bottom_offset=top_offset+int(crop_ratio*pic_height)\n          right_offset=left_offset+int(crop_ratio*pic_width)\n          f=cv2.resize(f[top_offset:bottom_offset,left_offset:right_offset,:],(input_height,input_width))\n        else:\n          f=f.resize((input_width, input_height))\n          f=np.asarray(f.convert('RGB'),dtype=np.uint8)          \n        x.append(f)\n      \n      \n      if random_crop and is_train:\n        y.append(filenames[i][1]-np.log(crop_ratio))\n      else:\n        y.append(filenames[i][1])\n      \n      count+=1\n      if count==batch_size:\n        x=np.array(x, dtype=np.float32)\n        y=np.array(y, dtype=np.float32)\n\n        inputs=x/255\n        targets=y       \n        x=[]\n        y=[]\n        count=0\n        yield inputs, targets\n\n\n\ndef aggregation_block(x_shallow, x_deep, deep_ch, out_ch):\n  x_deep= Conv2DTranspose(deep_ch, kernel_size=2, strides=2, padding='same', use_bias=False)(x_deep)\n  x_deep = BatchNormalization()(x_deep)   \n  x_deep = LeakyReLU(alpha=0.1)(x_deep)\n  x = Concatenate()([x_shallow, x_deep])\n  x=Conv2D(out_ch, kernel_size=1, strides=1, padding=\"same\")(x)\n  x = BatchNormalization()(x)   \n  x = LeakyReLU(alpha=0.1)(x)\n  return x\n  \n\n\ndef cbr(x, out_layer, kernel, stride):\n  x=Conv2D(out_layer, kernel_size=kernel, strides=stride, padding=\"same\")(x)\n  x = BatchNormalization()(x)\n  x = LeakyReLU(alpha=0.1)(x)\n  return x\n\ndef resblock(x_in,layer_n):\n  x=cbr(x_in,layer_n,3,1)\n  x=cbr(x,layer_n,3,1)\n  x=Add()([x,x_in])\n  return x  \n\n\n#I use the same network at CenterNet\ndef create_model(input_shape, size_detection_mode=True, aggregation=True):\n    input_layer = Input(input_shape)\n    \n    #resized input\n    input_layer_1=AveragePooling2D(2)(input_layer)\n    input_layer_2=AveragePooling2D(2)(input_layer_1)\n\n    #### ENCODER ####\n\n    x_0= cbr(input_layer, 16, 3, 2)#512-&gt;256\n    concat_1 = Concatenate()([x_0, input_layer_1])\n\n    x_1= cbr(concat_1, 32, 3, 2)#256-&gt;128\n    concat_2 = Concatenate()([x_1, input_layer_2])\n\n    x_2= cbr(concat_2, 64, 3, 2)#128-&gt;64\n    \n    x=cbr(x_2,64,3,1)\n    x=resblock(x,64)\n    x=resblock(x,64)\n    \n    x_3= cbr(x, 128, 3, 2)#64-&gt;32\n    x= cbr(x_3, 128, 3, 1)\n    x=resblock(x,128)\n    x=resblock(x,128)\n    x=resblock(x,128)\n    \n    x_4= cbr(x, 256, 3, 2)#32-&gt;16\n    x= cbr(x_4, 256, 3, 1)\n    x=resblock(x,256)\n    x=resblock(x,256)\n    x=resblock(x,256)\n    x=resblock(x,256)\n    x=resblock(x,256)\n \n    x_5= cbr(x, 512, 3, 2)#16-&gt;8\n    x= cbr(x_5, 512, 3, 1)\n    \n    x=resblock(x,512)\n    x=resblock(x,512)\n    x=resblock(x,512)\n    \n    if size_detection_mode:\n      x=GlobalAveragePooling2D()(x)\n      x=Dropout(0.2)(x)\n      out=Dense(1,activation=\"linear\")(x)\n    \n    else:#centernet mode\n    #### DECODER ####\n      x_1= cbr(x_1, output_layer_n, 1, 1)\n      x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n      x_2= cbr(x_2, output_layer_n, 1, 1)\n      x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n      x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n      x_3= cbr(x_3, output_layer_n, 1, 1)\n      x_3 = aggregation_block(x_3, x_4, output_layer_n, output_layer_n) \n      x_2 = aggregation_block(x_2, x_3, output_layer_n, output_layer_n)\n      x_1 = aggregation_block(x_1, x_2, output_layer_n, output_layer_n)\n      \n      x_4= cbr(x_4, output_layer_n, 1, 1)\n\n      x=cbr(x, output_layer_n, 1, 1)\n      x= UpSampling2D(size=(2, 2))(x)#8-&gt;16 tconvのがいいか\n\n      x = Concatenate()([x, x_4])\n      x=cbr(x, output_layer_n, 3, 1)\n      x= UpSampling2D(size=(2, 2))(x)#16-&gt;32\n    \n      x = Concatenate()([x, x_3])\n      x=cbr(x, output_layer_n, 3, 1)\n      x= UpSampling2D(size=(2, 2))(x)#32-&gt;64   128のがいいかも？ \n    \n      x = Concatenate()([x, x_2])\n      x=cbr(x, output_layer_n, 3, 1)\n      x= UpSampling2D(size=(2, 2))(x)#64-&gt;128 \n      \n      x = Concatenate()([x, x_1])\n      x=Conv2D(output_layer_n, kernel_size=3, strides=1, padding=\"same\")(x)\n      out = Activation(\"sigmoid\")(x)\n    \n    model=Model(input_layer, out)\n    \n    return model\n  \n    \n\ndef model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size=32):\n    hist = model.fit_generator(\n        Datagen_sizecheck_model(train_list,batch_size, is_train=True,random_crop=True),\n        steps_per_epoch = len(train_list) // batch_size,\n        epochs = n_epoch,\n        validation_data=Datagen_sizecheck_model(cv_list,batch_size, is_train=False,random_crop=False),\n        validation_steps = len(cv_list) // batch_size,\n        callbacks = [lr_schedule, model_checkpoint],#[early_stopping, reduce_lr, model_checkpoint],\n        shuffle = True,\n        verbose = 1\n    )\n    return hist\n\ndef lrs(epoch):\n    lr = 0.0005\n    if epoch&gt;10:\n        lr = 0.0001\n    return lr\n\nif __name__ == '__main__':\n        \n    parser = argparse.ArgumentParser()\n\n    parser.add_argument('--epochs', type=int, default=10)\n    parser.add_argument('--learning-rate', type=float, default=0.01)\n    parser.add_argument('--batch-size', type=int, default=128)\n    parser.add_argument('--gpu-count', type=int, default=os.environ['SM_NUM_GPUS'])\n    parser.add_argument('--model-dir', type=str, default=os.environ['SM_MODEL_DIR'])\n    parser.add_argument('--training', type=str, default=os.environ['SM_CHANNEL_TRAINING'])\n    parser.add_argument('--validation', type=str, default=os.environ['SM_CHANNEL_VALIDATION'])\n    \n    args, _ = parser.parse_known_args()\n    s3 = boto3.client('s3')\n\n    epochs     = args.epochs\n    lr         = args.learning_rate\n    batch_size = args.batch_size\n    gpu_count  = args.gpu_count\n    model_dir  = args.model_dir\n    training_dir   = args.training\n    validation_dir = args.validation\n\n    path_1= \"/opt/ml/input/data/training/train_csv/train.csv\"\n    path_2=\"/opt/ml/input/data/training/\"\n    path_3=\"/opt/ml/input/data/validation/\"\n    path_4=\"/opt/ml/input/data/validation/submission/sample_submission.csv\"\n    final_weights_step1 = \"/opt/ml/input/data/training/final_weights/final_weights_step1.hdf5\"\n\n\n    df_train=pd.read_csv(path_1)\n    #print(df_train.head())\n    #print(df_train.shape)\n    df_train=df_train.dropna(axis=0, how='any')#you can use nan data(page with no letter)\n    df_train=df_train.reset_index(drop=True)\n    #print(df_train.shape)\n\n    annotation_list_train=[]\n    category_names=set()\n\n    for i in range(len(df_train)):\n        ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,x,y,width,height for each picture\n        category_names=category_names.union({i for i in ann[:,0]})\n\n    category_names=sorted(category_names)\n    dict_cat={list(category_names)[j]:str(j) for j in range(len(category_names))}\n    inv_dict_cat={str(j):list(category_names)[j] for j in range(len(category_names))}\n    #print(dict_cat)\n    \n    for i in range(len(df_train)):\n        ann=np.array(df_train.loc[i,\"labels\"].split(\" \")).reshape(-1,5)#cat,left,top,width,height for each picture\n        for j,category_name in enumerate(ann[:,0]):\n            ann[j,0]=int(dict_cat[category_name])  \n        ann=ann.astype('int32')\n        ann[:,1]+=ann[:,3]//2#center_x\n        ann[:,2]+=ann[:,4]//2#center_y\n        annotation_list_train.append([\"{}{}.jpg\".format(path_2,df_train.loc[i,\"image_id\"]),ann])\n\n    df_submission=pd.read_csv(path_4)\n    id_test=path_3+df_submission[\"image_id\"]+\".jpg\"\n\n    aspect_ratio_pic_all=[]\n    aspect_ratio_pic_all_test=[]\n    average_letter_size_all=[]\n    train_input_for_size_estimate=[]\n    resize_dir=\"resized/\"\n    if os.path.exists(resize_dir) == False:os.mkdir(resize_dir)\n    for i in range(len(annotation_list_train)):\n        with Image.open(annotation_list_train[i][0]) as f:\n            width,height=f.size\n            area=width*height\n            aspect_ratio_pic=height/width\n            aspect_ratio_pic_all.append(aspect_ratio_pic)\n            letter_size=annotation_list_train[i][1][:,3]*annotation_list_train[i][1][:,4]\n            letter_size_ratio=letter_size/area\n        \n            average_letter_size=np.mean(letter_size_ratio)\n            average_letter_size_all.append(average_letter_size)\n            train_input_for_size_estimate.append([annotation_list_train[i][0],np.log(average_letter_size)])#logにしとく\n        \n\n    for i in range(len(id_test)):\n        with Image.open(id_test[i]) as f:\n            width,height=f.size\n            aspect_ratio_pic=height/width\n            aspect_ratio_pic_all_test.append(aspect_ratio_pic)\n\n    \n    K.clear_session()\n    model=create_model(input_shape=(input_height,input_width,3),size_detection_mode=True)\n\n    lr_schedule = LearningRateScheduler(lrs)\n    model_checkpoint = ModelCheckpoint(final_weights_step1, monitor = 'val_loss', verbose = 1,\n                                        save_best_only = True, save_weights_only = True, period = 1)\n    print(model.summary())\n\n    train_list, cv_list = train_test_split(train_input_for_size_estimate, random_state = 111,test_size = 0.2)\n\n\n    learning_rate=0.0005\n    n_epoch=10\n    batch_size=32\n\n    model.compile(loss=mean_squared_error, optimizer=Adam(lr=learning_rate))\n    hist = model_fit_sizecheck_model(model,train_list,cv_list,n_epoch,batch_size)\n\n    #model.save_weights('final_weights_step1.h5')\n    #model.load_weights(final_weights_step1)\n\n==========================================\n\nHere's the code for the SageMaker notebook file:\n\nimport os\nimport keras\nimport numpy as np\nimport sagemaker\nfrom sagemaker.tensorflow import TensorFlow\n\nsess = sagemaker.Session()\nrole = \"arn:aws:iam::xxxxxxxxxxxxxxxxx\"\n\ntraining_input_path = \"s3://xxxxxxxxxxxx/sagemaker-train-data\"\nvalidation_input_path = \"s3://xxxxxxxxxxxxx/sagemaker-test-data\"\n\ntf_estimator = TensorFlow(entry_point='sagemaker_centernet_script.py', \n                          role=role,\n                          train_instance_count=1, \n                          train_instance_type='ml.p3.2xlarge',\n                          framework_version='1.12', \n                          py_version='py3',\n                          script_mode=True,\n                          hyperparameters={\n                              'epochs': 20,\n                              'batch-size': 256,\n                              'learning-rate': 0.01}\n                         )\n\ntf_estimator.fit({'training': training_input_path, 'validation': validation_input_path})",
    "605068": "The slowdown is mostly explainable to the data generators not using keras.utils.Sequence, which when used enables multithreaded batch generation.  With the right code, I get avg 138 s/epoch at 2 s/step with workers=8 and max\\_queue\\_size=16. Note with Juypter and python use\\_multiprocessing should always be false due to a bug (unless you do some sleuthing to find some code that properly circumvents the known issue).  Additional note is that Step 1 cannot be too accurate as I've found that it breaks the heatmap into detecting everything on a page.",
    "605071": "Thanks, that is very interesting. But I was using the same exact settings in my CPU notebook as with my SageMaker notebook, which were also both the same as the original Kaggle notebook. So for comparison purposes like I was doing, would those settings you mentioned have mattered?",
    "605076": "So far from my experience, yes it would as setting up multithreading and having a good balance of threads/workers to gpu compute speed is shaving hours off on training but I am not sure on CPU only.  I am using my personal desktop with a Threadripper 1900X and a 1080 Ti for my numbers.  Surely with the SageMaker gpu instance you've used there is an extreme benefit.  I currently cannot test a CPU only config, but I think there is at least minimal benefit."
  },
  "source": "meta"
}