Skip to content
This repository was archived by the owner on Nov 2, 2023. It is now read-only.

Task B.4

Lucas Berry edited this page Sep 24, 2023 · 3 revisions

Machine Learning 1

create_model Function Changes

The create_model function was seperated into two functions, create_model; where the model's arcitecture is made and train_model; where the model is trained. This was done to allow for the model creation to be streamlined and to allow for running multiple models in the future. The create_model function now only creates the model and returns it. The train_model function takes the model and trains it with the given parameters. The train_model function also returns the trained model.

New create_model Function

def create_model(sequence_length, n_features=1, cell=LSTM, layer_size=[50, 50, 50], dropout=0.2, optimizer='adam', loss='mean_squared_error'):
    # Builds the model
    # Params:
    #   sequence_length (int)   : The historical sequence length used to predict
    #   n_features      (int)   : The number of features used to predict, default is 1
    #   cell            (func)  : The type of cell to use in the model, default is LSTM
    #   layer_size      (list)  : The size of each layer in the model, default is [50, 50, 50]
    #   dropout         (float) : The dropout rate to use in the model, default is 0.2
    #   optimizer       (str)   : The optimizer to use in the model, default is 'adam'
    #   loss            (str)   : The loss function to use in the model, default is 'mean_squared_error'
    # Returns:
    #   model           (model) : The model to be trained

    # Basic neural network
    model = Sequential()
    # Add layers to model based on the length of layer_size
    for i in range(len(layer_size)):
        # Create cell based on cell type
        if i == 0:
            # Add input layer that needs input shape defined
            model.add(InputLayer(input_shape=(sequence_length, n_features)))
            model.add(cell(units=layer_size[i], return_sequences=True))
            # For som eadvances explanation of return_sequences:
            # https://machinelearningmastery.com/return-sequences-and-return-states-for-lstms-in-keras/
            # https://www.dlology.com/blog/how-to-use-return_state-or-return_sequences-in-keras/
            # As explained there, for a stacked LSTM, you must set return_sequences=True 
            # when stacking LSTM layers so that the next LSTM layer has a 
            # three-dimensional sequence input. 
        elif i == len(layer_size) - 1:
            # Add output layer that doesn't need a return sequence
            model.add(cell(units=layer_size[i]))
        else:
            # Add hidden layer that needs a return sequence as it's a stacked LSTM but no input shape
            model.add(cell(units=layer_size[i], return_sequences=True))
        # Add dropout after each layer to prevent overfitting
        model.add(Dropout(dropout))
        # The Dropout layer randomly sets input units to 0 with a frequency of 
        # rate (= 0.2 above) at each step during training time, which helps 
        # prevent overfitting (one of the major problems of ML).
    # Add final dense layer to output prediction
    model.add(Dense(1))
    # Compile model with given optimizer and loss function
    model.compile(optimizer=optimizer, loss=loss)
    # The optimizer and loss are two important parameters when building an 
    # ANN model. Choosing a different optimizer/loss can affect the prediction
    # quality significantly. You should try other settings to learn; e.g.
    # optimizer='rmsprop'/'sgd'/'adadelta'/...
    # loss='mean_absolute_error'/'huber_loss'/'cosine_similarity'/...
    
    return model

The create_model function takes the sequence length, number of features, cell type, layer size, dropout rate, optimizer and loss function as parameters. The function then creates a sequential model and adds layers to it based on the layer size. The first layer is an input layer that needs the input shape defined. The input shape is the sequence length and number of features. The next layers are hidden layers that need a return sequence as it's a stacked LSTM but no input shape. The final layer is an output layer that doesn't need a return sequence. A dropout layer is added after each layer to prevent overfitting. The dropout rate is the rate at which the dropout layer randomly sets input units to 0 with a frequency of rate at each step during training time. The final layer is a dense layer that outputs the prediction. The model is then compiled with the given optimizer and loss function. The model is then returned.

train_model Function

def train_model(company, x_train, y_train, hyperparameters, refresh=True, save=True, model_dir='model'):
    # Trains the model
    # Params:
    #   company     (str)   : The company you want to train on, examples include AAPL, TESL, etc.
    #   x_train     (list)  : The x training data
    #   y_train     (list)  : The y training data
    #   refresh     (bool)  : Whether to retrain the model even if it exists, default is False
    #   save        (bool)  : Whether to save the model locally if it doesn't already exist, default is True
    #   model_dir   (str)   : Directory to store model, default is 'model'
    # Returns:
    #   model       (model) : The trained model
    
    # Creates model directory if it doesn't exist
    if not os.path.isdir(model_dir):
        os.mkdir(model_dir)

    # Model filename to ensure model is unique
    model_file_path = os.path.join(model_dir, f"model_{company}_{hyperparameters['sequence_length']}_{hyperparameters['cell'].__name__}_{'_'.join(map(str, hyperparameters['layer_size']))}_{hyperparameters['dropout']}_{hyperparameters['optimizer']}_{hyperparameters['loss']}.keras")

    # Checks if data file with same data exists
    if os.path.exists(model_file_path) and not refresh:
        # If file exists and data shouldn't be updated, import as pandas data frame object
        # 'index_col=0' makes the date the index rather than making a new coloumn
        model = load_model(model_file_path)
        return model

    model = create_model(hyperparameters['sequence_length'], hyperparameters['n_features'], hyperparameters['cell'], hyperparameters['layer_size'], hyperparameters['dropout'], hyperparameters['optimizer'], hyperparameters['loss'])

    # Now we are going to train this model with our training data 
    # (x_train, y_train)
    model.fit(x_train, y_train, epochs=25, batch_size=32)
    # Other parameters to consider: How many rounds(epochs) are we going to 
    # train our model? Typically, the more the better, but be careful about
    # overfitting!
    # What about batch_size? Well, again, please refer to 
    # Lecture Week 6 (COS30018): If you update your model for each and every 
    # input sample, then there are potentially 2 issues: 1. If you training 
    # data is very big (billions of input samples) then it will take VERY long;
    # 2. Each and every input can immediately makes changes to your model
    # (a souce of overfitting). Thus, we do this in batches: We'll look at
    # the aggreated errors/losses from a batch of, say, 32 input samples
    # and update our model based on this aggregated loss.

    if save:
        model.save(model_file_path)

    return model

The train_model function takes the company, x training data, y training data, hyperparameters, refresh, save and model directory as parameters. The function then creates the model directory if it doesn't exist. The model filename is then created to ensure the model is unique. The function then checks if the model file exists and if the data shouldn't be updated. If the file exists and the data shouldn't be updated, the model is loaded and returned. The function then creates the model with the given hyperparameters. The model is then trained with the x and y training data. The model is then saved if save is true. The model is then returned.

Testing Other Model Algorithms

The default algorithm for the model is LSTM. LSTM is also used by P1 and was pretty good with it's predictions. The other algorithms that were tested were GRU and RNN. The results of the testing can be seen below. A function that calculates the average error was also created to make it easier to compare the models.

def error(test_df, predicted_df, feature_columns=['Open', 'High', 'Low', 'Close']):
    # Uses model and data to assess prediction
    # Params:
    #   test_df         (df)    : The test data downloaded from yahoo
    #   predicted_df    (df)    : The predictions based on the test data
    #   feature_columns (list)  : The list of features graphed from the prediction, default is everything grabbed from yahoo
    # Returns:
    #   feature_error   (dict)  : The average error and average error percentage for each feature

    feature_error = {}
    for column in feature_columns:
        # Calculate the average error for each feature
        average_error = np.mean(np.abs(test_df[column] - predicted_df[column]))
        # Calculate the average error percentage for each feature
        average_error_percentage = average_error / np.mean(test_df[column])
        # Add the average error and average error percentage to the feature_error dict
        feature_error[column] = (average_error, average_error_percentage)

    return feature_error

LSTM

image

image

GRU

image

image

RNN

image

image

Results

The function for finding average error is far from perfect but it does give a good idea of how the models compare. All else being equal, It shows that LSTM and RNN are very similar with GRU being considerably more acurate, at least in the current implementation. In time it would be desireable to set up the model to train based on each feature column rather than just 'Close' as this would allow for more accurate predictions. It would also be desireable to test more hyperparameters to find the best ones for each model.

Clone this wiki locally