{
  "id": 76361,
  "title": "why it will stop at the 3 epoch during the training while using keras",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76361",
  "author_name": "",
  "post_date": "2019-01-01T20:07:28.770382900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>here is the training info </p>\n\n<p>Train on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 513s 393us/step - loss: 0.1122 - acc: 0.9557 - val_loss: 0.0997 - val_acc: 0.9625\n1307/1307 [==============================] - 0s 307us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 509s 390us/step - loss: 0.0987 - acc: 0.9606 - val_loss: 0.0964 - val_acc: 0.9579\n1307/1307 [==============================] - 0s 130us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n 471552/1304815 [=========&gt;....................] - ETA: 5:24 - loss: 0.0922 - acc: 0.9628</p>\n\n<p>as u can see, it always stop at the third epoch without any error at the training step of first k-fold dataset.  Because of this , my kernel can never finish the whole k-fold groups.</p>\n\n<p>Could anyone tell me why??</p>",
  "messages": [
    {
      "id": "448669",
      "postDate": "01/01/2019 20:07:28",
      "content": "<p>here is the training info </p>\n\n<p>Train on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 513s 393us/step - loss: 0.1122 - acc: 0.9557 - val_loss: 0.0997 - val_acc: 0.9625\n1307/1307 [==============================] - 0s 307us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 509s 390us/step - loss: 0.0987 - acc: 0.9606 - val_loss: 0.0964 - val_acc: 0.9579\n1307/1307 [==============================] - 0s 130us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n 471552/1304815 [=========&gt;....................] - ETA: 5:24 - loss: 0.0922 - acc: 0.9628</p>\n\n<p>as u can see, it always stop at the third epoch without any error at the training step of first k-fold dataset.  Because of this , my kernel can never finish the whole k-fold groups.</p>\n\n<p>Could anyone tell me why??</p>",
      "rawMarkdown": "here is the training info \n\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 513s 393us/step - loss: 0.1122 - acc: 0.9557 - val_loss: 0.0997 - val_acc: 0.9625\n1307/1307 [==============================] - 0s 307us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 509s 390us/step - loss: 0.0987 - acc: 0.9606 - val_loss: 0.0964 - val_acc: 0.9579\n1307/1307 [==============================] - 0s 130us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n 471552/1304815 [=========&gt;....................] - ETA: 5:24 - loss: 0.0922 - acc: 0.9628\n\nas u can see, it always stop at the third epoch without any error at the training step of first k-fold dataset.  Because of this , my kernel can never finish the whole k-fold groups.\n\nCould anyone tell me why??",
      "votes": null
    },
    {
      "id": "448677",
      "postDate": "01/01/2019 20:28:29",
      "content": "<p>Hi! asvgdsf, hard to tell what is the problem from just the training output, can you share more information, for example the block of code for Keras Model.</p>",
      "rawMarkdown": "Hi! asvgdsf, hard to tell what is the problem from just the training output, can you share more information, for example the block of code for Keras Model.",
      "votes": null
    },
    {
      "id": "448680",
      "postDate": "01/01/2019 20:39:11",
      "content": "<p>```</p>\n\n<h1>code block</h1>\n\n<p>def model_capsule(embedding_matrix):\n    K.clear_session() <br>\n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(rate=0.2)(x)\n    x = Bidirectional(CuDNNGRU(100, return_sequences=True, \n                                kernel_initializer=glorot_normal(seed=12300), recurrent_initializer=orthogonal(gain=1.0, seed=10000)))(x)\n    #print(x.shape,'---') 72  200\n    x = Capsule(num_capsule=10, dim_capsule=10, routings=4, share_weights=True)(x)\n    #print(x.shape,'-----') 10  10\n    x = Flatten()(x)\n    #print(x.shape,'-----') 100\n    x = Dense(100, activation=\"relu\", kernel_initializer=glorot_normal(seed=12300))(x)\n    x = Dropout(0.12)(x)\n    x = BatchNormalization()(x)\n    x = Dense(1, activation=\"sigmoid\")(x)\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    return model\n```</p>\n\n<p>```</p>\n\n<h1>code block</h1>\n\n<p>bestthr = []\ntest_y = []\nsplits = list(StratifiedKFold(n_splits=10, shuffle=True, random_state=1212).split(train_X, train_y))\nfor i, (train_index, valid_index) in enumerate(splits):\n    X_train, X_val, Y_train, Y_val = train_X[train_index], train_X[valid_index], train_y[train_index],     train_y[valid_index]\n    try:\n        cap_model = model_capsule(embedding_matrix)\n        cap_model.fit(X_train, Y_train, batch_size=2048, epochs=3, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cap_pred_val_y = cap_model.predict([X_val], batch_size=1024, verbose=1)\n    cap_pred_test_y = cap_model.predict([test_X], batch_size=1024, verbose=1)\n    try:\n        cnn_model = model_cnn(embedding_matrix_1)\n        cnn_model.fit(X_train, Y_train, batch_size=2048, epochs=2, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cnn_pred_val_y = cnn_model.predict([X_val], batch_size=10240, verbose=1)\n    cnn_pred_test_y = cnn_model.predict([test_X], batch_size=10240, verbose=1)\n    coefs,max_t = get_conf(cap_pred_val_y,cnn_pred_val_y,Y_val)\n    bestthr.append(max_t)\n    tmp_test = coefs[0]*cap_pred_test_y + coefs[1]*cnn_pred_test_y\n    test_y.append(tmp_test)\n```</p>",
      "rawMarkdown": "```\n# code block\ndef model_capsule(embedding_matrix):\n    K.clear_session()       \n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(rate=0.2)(x)\n    x = Bidirectional(CuDNNGRU(100, return_sequences=True, \n                                kernel_initializer=glorot_normal(seed=12300), recurrent_initializer=orthogonal(gain=1.0, seed=10000)))(x)\n    #print(x.shape,'---') 72  200\n    x = Capsule(num_capsule=10, dim_capsule=10, routings=4, share_weights=True)(x)\n    #print(x.shape,'-----') 10  10\n    x = Flatten()(x)\n    #print(x.shape,'-----') 100\n    x = Dense(100, activation=\"relu\", kernel_initializer=glorot_normal(seed=12300))(x)\n    x = Dropout(0.12)(x)\n    x = BatchNormalization()(x)\n    x = Dense(1, activation=\"sigmoid\")(x)\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    return model\n```\n\n\n\n\n```\n# code block\nbestthr = []\ntest_y = []\nsplits = list(StratifiedKFold(n_splits=10, shuffle=True, random_state=1212).split(train_X, train_y))\nfor i, (train_index, valid_index) in enumerate(splits):\n    X_train, X_val, Y_train, Y_val = train_X[train_index], train_X[valid_index], train_y[train_index],     train_y[valid_index]\n    try:\n        cap_model = model_capsule(embedding_matrix)\n        cap_model.fit(X_train, Y_train, batch_size=2048, epochs=3, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cap_pred_val_y = cap_model.predict([X_val], batch_size=1024, verbose=1)\n    cap_pred_test_y = cap_model.predict([test_X], batch_size=1024, verbose=1)\n    try:\n        cnn_model = model_cnn(embedding_matrix_1)\n        cnn_model.fit(X_train, Y_train, batch_size=2048, epochs=2, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cnn_pred_val_y = cnn_model.predict([X_val], batch_size=10240, verbose=1)\n    cnn_pred_test_y = cnn_model.predict([test_X], batch_size=10240, verbose=1)\n    coefs,max_t = get_conf(cap_pred_val_y,cnn_pred_val_y,Y_val)\n    bestthr.append(max_t)\n    tmp_test = coefs[0]*cap_pred_test_y + coefs[1]*cnn_pred_test_y\n    test_y.append(tmp_test)\n```",
      "votes": null
    },
    {
      "id": "448682",
      "postDate": "01/01/2019 20:44:30",
      "content": "<p>There's a known bug limiting code cell outputs to a certain number of characters. </p>\n\n<p>See: <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76092\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76092</a></p>",
      "rawMarkdown": "There's a known bug limiting code cell outputs to a certain number of characters. \n\nSee: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76092",
      "votes": null
    },
    {
      "id": "448684",
      "postDate": "01/01/2019 20:50:24",
      "content": "<p>here is the main part, the others are from the public kernels, and i find that all the commit kernels will be interrupted at the third epoch  have u check your commit kernel?</p>",
      "rawMarkdown": "here is the main part, the others are from the public kernels, and i find that all the commit kernels will be interrupted at the third epoch  have u check your commit kernel?",
      "votes": null
    },
    {
      "id": "448685",
      "postDate": "01/01/2019 20:54:35",
      "content": "<p>thanks Bilal</p>",
      "rawMarkdown": "thanks Bilal",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 448677,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "01/01/2019 20:28:29",
      "content": "<p>Hi! asvgdsf, hard to tell what is the problem from just the training output, can you share more information, for example the block of code for Keras Model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 448680,
          "author_name": "sdafweagvdsv",
          "author_url": "",
          "post_date": "01/01/2019 20:39:11",
          "content": "<p>```</p>\n\n<h1>code block</h1>\n\n<p>def model_capsule(embedding_matrix):\n    K.clear_session() <br>\n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(rate=0.2)(x)\n    x = Bidirectional(CuDNNGRU(100, return_sequences=True, \n                                kernel_initializer=glorot_normal(seed=12300), recurrent_initializer=orthogonal(gain=1.0, seed=10000)))(x)\n    #print(x.shape,'---') 72  200\n    x = Capsule(num_capsule=10, dim_capsule=10, routings=4, share_weights=True)(x)\n    #print(x.shape,'-----') 10  10\n    x = Flatten()(x)\n    #print(x.shape,'-----') 100\n    x = Dense(100, activation=\"relu\", kernel_initializer=glorot_normal(seed=12300))(x)\n    x = Dropout(0.12)(x)\n    x = BatchNormalization()(x)\n    x = Dense(1, activation=\"sigmoid\")(x)\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    return model\n```</p>\n\n<p>```</p>\n\n<h1>code block</h1>\n\n<p>bestthr = []\ntest_y = []\nsplits = list(StratifiedKFold(n_splits=10, shuffle=True, random_state=1212).split(train_X, train_y))\nfor i, (train_index, valid_index) in enumerate(splits):\n    X_train, X_val, Y_train, Y_val = train_X[train_index], train_X[valid_index], train_y[train_index],     train_y[valid_index]\n    try:\n        cap_model = model_capsule(embedding_matrix)\n        cap_model.fit(X_train, Y_train, batch_size=2048, epochs=3, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cap_pred_val_y = cap_model.predict([X_val], batch_size=1024, verbose=1)\n    cap_pred_test_y = cap_model.predict([test_X], batch_size=1024, verbose=1)\n    try:\n        cnn_model = model_cnn(embedding_matrix_1)\n        cnn_model.fit(X_train, Y_train, batch_size=2048, epochs=2, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cnn_pred_val_y = cnn_model.predict([X_val], batch_size=10240, verbose=1)\n    cnn_pred_test_y = cnn_model.predict([test_X], batch_size=10240, verbose=1)\n    coefs,max_t = get_conf(cap_pred_val_y,cnn_pred_val_y,Y_val)\n    bestthr.append(max_t)\n    tmp_test = coefs[0]*cap_pred_test_y + coefs[1]*cnn_pred_test_y\n    test_y.append(tmp_test)\n```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448684,
          "author_name": "sdafweagvdsv",
          "author_url": "",
          "post_date": "01/01/2019 20:50:24",
          "content": "<p>here is the main part, the others are from the public kernels, and i find that all the commit kernels will be interrupted at the third epoch  have u check your commit kernel?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448682,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "01/01/2019 20:44:30",
      "content": "<p>There's a known bug limiting code cell outputs to a certain number of characters. </p>\n\n<p>See: <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76092\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76092</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 448685,
          "author_name": "sdafweagvdsv",
          "author_url": "",
          "post_date": "01/01/2019 20:54:35",
          "content": "<p>thanks Bilal</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "448669": "here is the training info \n\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 513s 393us/step - loss: 0.1122 - acc: 0.9557 - val_loss: 0.0997 - val_acc: 0.9625\n1307/1307 [==============================] - 0s 307us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n1304815/1304815 [==============================] - 509s 390us/step - loss: 0.0987 - acc: 0.9606 - val_loss: 0.0964 - val_acc: 0.9579\n1307/1307 [==============================] - 0s 130us/step\nTrain on 1304815 samples, validate on 1307 samples\nEpoch 1/1\n 471552/1304815 [=========&gt;....................] - ETA: 5:24 - loss: 0.0922 - acc: 0.9628\n\nas u can see, it always stop at the third epoch without any error at the training step of first k-fold dataset.  Because of this , my kernel can never finish the whole k-fold groups.\n\nCould anyone tell me why??",
    "448677": "Hi! asvgdsf, hard to tell what is the problem from just the training output, can you share more information, for example the block of code for Keras Model.",
    "448680": "```\n# code block\ndef model_capsule(embedding_matrix):\n    K.clear_session()       \n    inp = Input(shape=(maxlen,))\n    x = Embedding(max_features, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = SpatialDropout1D(rate=0.2)(x)\n    x = Bidirectional(CuDNNGRU(100, return_sequences=True, \n                                kernel_initializer=glorot_normal(seed=12300), recurrent_initializer=orthogonal(gain=1.0, seed=10000)))(x)\n    #print(x.shape,'---') 72  200\n    x = Capsule(num_capsule=10, dim_capsule=10, routings=4, share_weights=True)(x)\n    #print(x.shape,'-----') 10  10\n    x = Flatten()(x)\n    #print(x.shape,'-----') 100\n    x = Dense(100, activation=\"relu\", kernel_initializer=glorot_normal(seed=12300))(x)\n    x = Dropout(0.12)(x)\n    x = BatchNormalization()(x)\n    x = Dense(1, activation=\"sigmoid\")(x)\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    return model\n```\n\n\n\n\n```\n# code block\nbestthr = []\ntest_y = []\nsplits = list(StratifiedKFold(n_splits=10, shuffle=True, random_state=1212).split(train_X, train_y))\nfor i, (train_index, valid_index) in enumerate(splits):\n    X_train, X_val, Y_train, Y_val = train_X[train_index], train_X[valid_index], train_y[train_index],     train_y[valid_index]\n    try:\n        cap_model = model_capsule(embedding_matrix)\n        cap_model.fit(X_train, Y_train, batch_size=2048, epochs=3, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cap_pred_val_y = cap_model.predict([X_val], batch_size=1024, verbose=1)\n    cap_pred_test_y = cap_model.predict([test_X], batch_size=1024, verbose=1)\n    try:\n        cnn_model = model_cnn(embedding_matrix_1)\n        cnn_model.fit(X_train, Y_train, batch_size=2048, epochs=2, validation_data=(X_val, Y_val))\n    except:\n        pass\n    cnn_pred_val_y = cnn_model.predict([X_val], batch_size=10240, verbose=1)\n    cnn_pred_test_y = cnn_model.predict([test_X], batch_size=10240, verbose=1)\n    coefs,max_t = get_conf(cap_pred_val_y,cnn_pred_val_y,Y_val)\n    bestthr.append(max_t)\n    tmp_test = coefs[0]*cap_pred_test_y + coefs[1]*cnn_pred_test_y\n    test_y.append(tmp_test)\n```",
    "448682": "There's a known bug limiting code cell outputs to a certain number of characters. \n\nSee: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/76092",
    "448684": "here is the main part, the others are from the public kernels, and i find that all the commit kernels will be interrupted at the third epoch  have u check your commit kernel?",
    "448685": "thanks Bilal"
  },
  "source": "meta"
}