Repository navigation
Task B.4
The create_model function was seperated into two functions, create_model; where the model's arcitecture is made and train_model; where the model is trained. This was done to allow for the model creation to be streamlined and to allow for running multiple models in the future. The create_model function now only creates the model and returns it. The train_model function takes the model and trains it with the given parameters. The train_model function also returns the trained model.
def create_model(sequence_length, n_features=1, cell=LSTM, layer_size=[50, 50, 50], dropout=0.2, optimizer='adam', loss='mean_squared_error'):
# Builds the model
# Params:
# sequence_length (int) : The historical sequence length used to predict
# n_features (int) : The number of features used to predict, default is 1
# cell (func) : The type of cell to use in the model, default is LSTM
# layer_size (list) : The size of each layer in the model, default is [50, 50, 50]
# dropout (float) : The dropout rate to use in the model, default is 0.2
# optimizer (str) : The optimizer to use in the model, default is 'adam'
# loss (str) : The loss function to use in the model, default is 'mean_squared_error'
# Returns:
# model (model) : The model to be trained
# Basic neural network
model = Sequential()
# Add layers to model based on the length of layer_size
for i in range(len(layer_size)):
# Create cell based on cell type
if i == 0:
# Add input layer that needs input shape defined
model.add(InputLayer(input_shape=(sequence_length, n_features)))
model.add(cell(units=layer_size[i], return_sequences=True))
# For som eadvances explanation of return_sequences:
# https://machinelearningmastery.com/return-sequences-and-return-states-for-lstms-in-keras/
# https://www.dlology.com/blog/how-to-use-return_state-or-return_sequences-in-keras/
# As explained there, for a stacked LSTM, you must set return_sequences=True
# when stacking LSTM layers so that the next LSTM layer has a
# three-dimensional sequence input.
elif i == len(layer_size) - 1:
# Add output layer that doesn't need a return sequence
model.add(cell(units=layer_size[i]))
else:
# Add hidden layer that needs a return sequence as it's a stacked LSTM but no input shape
model.add(cell(units=layer_size[i], return_sequences=True))
# Add dropout after each layer to prevent overfitting
model.add(Dropout(dropout))
# The Dropout layer randomly sets input units to 0 with a frequency of
# rate (= 0.2 above) at each step during training time, which helps
# prevent overfitting (one of the major problems of ML).
# Add final dense layer to output prediction
model.add(Dense(1))
# Compile model with given optimizer and loss function
model.compile(optimizer=optimizer, loss=loss)
# The optimizer and loss are two important parameters when building an
# ANN model. Choosing a different optimizer/loss can affect the prediction
# quality significantly. You should try other settings to learn; e.g.
# optimizer='rmsprop'/'sgd'/'adadelta'/...
# loss='mean_absolute_error'/'huber_loss'/'cosine_similarity'/...
return modelThe create_model function takes the sequence length, number of features, cell type, layer size, dropout rate, optimizer and loss function as parameters. The function then creates a sequential model and adds layers to it based on the layer size. The first layer is an input layer that needs the input shape defined. The input shape is the sequence length and number of features. The next layers are hidden layers that need a return sequence as it's a stacked LSTM but no input shape. The final layer is an output layer that doesn't need a return sequence. A dropout layer is added after each layer to prevent overfitting. The dropout rate is the rate at which the dropout layer randomly sets input units to 0 with a frequency of rate at each step during training time. The final layer is a dense layer that outputs the prediction. The model is then compiled with the given optimizer and loss function. The model is then returned.
def train_model(company, x_train, y_train, hyperparameters, refresh=True, save=True, model_dir='model'):
# Trains the model
# Params:
# company (str) : The company you want to train on, examples include AAPL, TESL, etc.
# x_train (list) : The x training data
# y_train (list) : The y training data
# refresh (bool) : Whether to retrain the model even if it exists, default is False
# save (bool) : Whether to save the model locally if it doesn't already exist, default is True
# model_dir (str) : Directory to store model, default is 'model'
# Returns:
# model (model) : The trained model
# Creates model directory if it doesn't exist
if not os.path.isdir(model_dir):
os.mkdir(model_dir)
# Model filename to ensure model is unique
model_file_path = os.path.join(model_dir, f"model_{company}_{hyperparameters['sequence_length']}_{hyperparameters['cell'].__name__}_{'_'.join(map(str, hyperparameters['layer_size']))}_{hyperparameters['dropout']}_{hyperparameters['optimizer']}_{hyperparameters['loss']}.keras")
# Checks if data file with same data exists
if os.path.exists(model_file_path) and not refresh:
# If file exists and data shouldn't be updated, import as pandas data frame object
# 'index_col=0' makes the date the index rather than making a new coloumn
model = load_model(model_file_path)
return model
model = create_model(hyperparameters['sequence_length'], hyperparameters['n_features'], hyperparameters['cell'], hyperparameters['layer_size'], hyperparameters['dropout'], hyperparameters['optimizer'], hyperparameters['loss'])
# Now we are going to train this model with our training data
# (x_train, y_train)
model.fit(x_train, y_train, epochs=25, batch_size=32)
# Other parameters to consider: How many rounds(epochs) are we going to
# train our model? Typically, the more the better, but be careful about
# overfitting!
# What about batch_size? Well, again, please refer to
# Lecture Week 6 (COS30018): If you update your model for each and every
# input sample, then there are potentially 2 issues: 1. If you training
# data is very big (billions of input samples) then it will take VERY long;
# 2. Each and every input can immediately makes changes to your model
# (a souce of overfitting). Thus, we do this in batches: We'll look at
# the aggreated errors/losses from a batch of, say, 32 input samples
# and update our model based on this aggregated loss.
if save:
model.save(model_file_path)
return modelThe train_model function takes the company, x training data, y training data, hyperparameters, refresh, save and model directory as parameters. The function then creates the model directory if it doesn't exist. The model filename is then created to ensure the model is unique. The function then checks if the model file exists and if the data shouldn't be updated. If the file exists and the data shouldn't be updated, the model is loaded and returned. The function then creates the model with the given hyperparameters. The model is then trained with the x and y training data. The model is then saved if save is true. The model is then returned.
The default algorithm for the model is LSTM. LSTM is also used by P1 and was pretty good with it's predictions. The other algorithms that were tested were GRU and RNN. The results of the testing can be seen below. A function that calculates the average error was also created to make it easier to compare the models.
def error(test_df, predicted_df, feature_columns=['Open', 'High', 'Low', 'Close']):
# Uses model and data to assess prediction
# Params:
# test_df (df) : The test data downloaded from yahoo
# predicted_df (df) : The predictions based on the test data
# feature_columns (list) : The list of features graphed from the prediction, default is everything grabbed from yahoo
# Returns:
# feature_error (dict) : The average error and average error percentage for each feature
feature_error = {}
for column in feature_columns:
# Calculate the average error for each feature
average_error = np.mean(np.abs(test_df[column] - predicted_df[column]))
# Calculate the average error percentage for each feature
average_error_percentage = average_error / np.mean(test_df[column])
# Add the average error and average error percentage to the feature_error dict
feature_error[column] = (average_error, average_error_percentage)
return feature_error





The function for finding average error is far from perfect but it does give a good idea of how the models compare. All else being equal, It shows that LSTM and RNN are very similar with GRU being considerably more acurate, at least in the current implementation. In time it would be desireable to set up the model to train based on each feature column rather than just 'Close' as this would allow for more accurate predictions. It would also be desireable to test more hyperparameters to find the best ones for each model.